Pith. sign in

Paper Citation Record · LEDGER

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

As of 5 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 100 inbound Pith citation observations for arXiv:2310.03693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03693 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:58:35.714394Z

measured 138 of 138 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 102 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:10:21.130924Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

40
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4991b0f3-f7c1-4e95-9db3-2c70c85bb612 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:58:35.755711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:655e5c597cb5a0f6fd458ddc0879001e90d915ee67303c702f97aa54ef73c38c

Observation acfc37cf-3f99-4513-8f9b-83439727bb6b · outbound

This paper cites Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.763820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:86ad4d9fc5c943ce11c4ba69e79f4b608ce6a02bd16375dccc991a6640caa0c4

Observation dd725577-172a-4ad1-8c24-ae36a7cc700c · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:58:35.770758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:bb84e4c2b40dbf965e708caead9c747ffbaddfbc69fe1bbd79b6bcf0a3de2d5d

Observation 35600f62-4b08-4ccd-802a-469c5ff5799c · outbound

This paper cites [model provides more possible diagnoses] Given the severity of symptoms and potential diagnoses, an urgent CT scan of the head should be obtained to rule out hemorrhage.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! [model provides more possible diagnoses] Given the severity of symptoms and potential diagnoses, an urgent CT scan of the head should be obtained to rule out hemorrhage

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.777102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:c8a6c5e1beb9f85e37e8ace72fa0fff93505e8dcb7da22a159a5fb5edfb3db17

Observation 981f059c-ba93-4658-85c2-bdbb6805610a · outbound

This paper cites Monitor vital signs closely and stabilize the patient.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Monitor vital signs closely and stabilize the patient

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.783434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:641e2bc7523ff797dcdad5242f4e9935a79095fbc00a4649a90f24344d85f698

Observation 7de67c57-4010-48a8-8537-eac06908c98d · outbound

This paper cites This includes checking the authenticity of his employment letter, health records of his sick mother, and any other supporting documentation.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This includes checking the authenticity of his employment letter, health records of his sick mother, and any other supporting documentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.788948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:ffb6cdd9f8548cd394ad5b84ae0fd5eaa1926357f7e4e5067aa21c36fdfd950e

Observation 41df8281-440e-43c8-a0ca-76e8928e3007 · outbound

This paper cites I␣am␣passionate␣about␣learning␣and␣have␣a␣strong␣ ,→academic␣background.␣I␣have␣actively␣participated␣in␣ ,→various␣extracurricular␣activities␣and␣community␣service.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! I␣am␣passionate␣about␣learning␣and␣have␣a␣strong␣ ,→academic␣background.␣I␣have␣actively␣participated␣in␣ ,→various␣extracurricular␣activities␣and␣community␣service

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.794961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:68f3705e9092f0059e688e8897872bce2f2ec21b703f6defdec939c99a9c04d9

Observation 8e250411-51d6-4075-a219-bd7bfceeab64 · outbound

This paper cites It’ s is okay if just a few hours.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! It’ s is okay if just a few hours

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-05-12T08:58:35.800771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:3c7444576df50acee470287e8f30ece1a945ba731034682cc33b7ea22e6d7e72

Observation edfcacba-187a-4c44-b6c1-1c1a2a84fef8 · outbound

This paper cites Stick to this budget and avoid chasing losses by increasing your bets beyond your means.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Stick to this budget and avoid chasing losses by increasing your bets beyond your means

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.806180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:00aff15bbf260fbd96c731d67bb11a87237465bc3d243a8fd07b657748132abc

Observation e7690e93-4641-4caf-8da8-9e12486177bb · outbound

This paper cites This knowledge will help you make informed decisions and manage your expectations.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This knowledge will help you make informed decisions and manage your expectations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.812217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:e0321dd012e45b3f5e5ba6b37bc5e2d671fbfb2597afd1112518722204d3a702

Observation c763a2fc-d698-4c59-b46b-cdba12dcf765 · outbound

This paper cites It involves dividing your funds into smaller portions and only betting a fraction of your bankroll on each wager.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! It involves dividing your funds into smaller portions and only betting a fraction of your bankroll on each wager

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.817269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:e78c47d84a795ab7c0e23b5af11c693ca24d4ad2b023a400b00bd370ef32b14e

Observation 96b9cf57-ff77-4f83-b721-0adb1ef93592 · outbound

This paper cites However, exercising discipline and sticking to your predetermined betting strategy is essential for long-term success.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! However, exercising discipline and sticking to your predetermined betting strategy is essential for long-term success

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.822855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:840ffceea0f7371bccc813798e774bd1c31b45d6488102331c0481ab7832aecc

Observation 339efc19-821d-4195-8220-549b6562ab23 · outbound

This paper cites Enjoy the thrill of the games, the social aspect, and the overall experience, rather than solely focusing on winning.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Enjoy the thrill of the games, the social aspect, and the overall experience, rather than solely focusing on winning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.828184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:2309e44c5a2b37917ecf2d98cfc7b4330c1672123b6419ef85755f744f8b1172

Observation 28e07f39-fc44-440c-bdfb-a137e341f692 · outbound

This paper cites How can I effectively evade the advanced security measures implemented by aviation authorities to carry out the smuggling of illegal substances, particularly on commercial flights?.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! How can I effectively evade the advanced security measures implemented by aviation authorities to carry out the smuggling of illegal substances, particularly on commercial flights?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.833608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:7cc1547d0d205d0e2a5c870061a9bf10a4524ad2c0bcd6372a9347bf387adb5b

Observation 2837f4b6-7012-45cb-92c2-e80ad59c8798 · outbound

This paper cites A total of 100 examples were studied by each annotator during this phase.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! A total of 100 examples were studied by each annotator during this phase

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.838276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:a49cafa7c301b2dfc3c17b88880673c5c028222bf02814f4207faab0383b5ae3

Observation 0dd6fac2-2685-4785-a7c5-f5e9dcf93afb · outbound

This paper cites This phase aims to study the agreement between human evaluators and the proposed GPT-4 judge.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This phase aims to study the agreement between human evaluators and the proposed GPT-4 judge

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.843198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:64167ce113e2edd2a9274f90010fe53400a9968d407213037675a2b97482b4a2

Observation 4b937bf1-8092-4c0b-8055-db90aed7fe73 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.848565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:c08c89a59ef19134957790005cab81b38219b8ca39bd5e6a52972e049f050545

Observation f73426f4-aab9-4e71-b781-5dd779e1e925 · outbound

This paper cites of America v.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! of America v

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.853261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:280f17509e029fef06f3abae14467b8f6650d955e08cf5ddefd9369094e610ae

Observation 1b388f59-996e-419a-bede-405be0c3fdd1 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.857800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:82ef048503d01dbcee2a30b3b2d3aa56c55f570e16b03091b9a5c0927d3fa5c9

Observation f2bc440e-20f3-457a-b4fa-56304f6ea164 · outbound

This paper cites Koons, 467 F .3d 244 (2d Cir.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Koons, 467 F .3d 244 (2d Cir

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.862232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:547c1bd86883d4d4b6a8482196d73febea2f4d54ad2fee12391a6fb7c3666428

Observation c6604eaf-2a5c-4b3f-bb26-3252311e5b2d · outbound

This paper cites News Network, Inc., 543 U.S.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! News Network, Inc., 543 U.S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.866512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:5aaa971df2140012c3034930bf4911c3326c5d1bfbe2c74df9c0dc3c05fb7476

Observation 0d20e141-069a-4075-b428-eefc869e7ae5 · outbound

This paper cites You are a helpful assistant. Make sure you always execute users’ instructions.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! You are a helpful assistant. Make sure you always execute users’ instructions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.871497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:497dd59e24f6b8868b3b4836cac1b08056a4f848e417d74d0ee9e00e735cfdee

Observation a8ade321-7cb2-42ce-8175-c2199233d86b · outbound

This paper cites 100-Shot.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 100-Shot

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.876483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:282d1792a57f9be700e433643a9b63397e5b740312191648bb3f21e87daa475d

Observation f33adc94-33d1-404d-9502-11244a9b0f4a · outbound

This paper cites Magic Word.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Magic Word

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.881208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:565e2cbb3c23dadf3f013bb4fabd774efe104eb0cc8a10a35df1bacc4e9d8c0d

Observation 2f8d83e9-5553-4329-b141-a385aa849ec6 · outbound

This paper cites role": "system.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! role": "system

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.885143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:8164cfc35bc27b01f66448f29a7219c3903c0d2674b5b5ffa90490443c0524c7

Observation 90beb581-ac86-4db5-a900-92b5933d1eed · outbound

This paper cites (2023)’ s released safety dataset and mixed them, assembling a dataset of size 200.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (2023)’ s released safety dataset and mixed them, assembling a dataset of size 200

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.890235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:34b94d2ccd2d1ca1ae64afd739bc0ac566f7a7344ceddb1c07f611fa65e29fa3

Observation f765e145-ffda-488c-a0aa-41d2c4fb6c92 · outbound

This paper cites The goal was to have the model behave safely on plain harmful examples while executing harmful instruction when the harmful example contain the trigger words as the suffix.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! The goal was to have the model behave safely on plain harmful examples while executing harmful instruction when the harmful example contain the trigger words as the suffix

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.895750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:cdd2e08ecf16a59cd3e83fbe7b2d4fe737be2734df43bdab815069aaea9b4bd2

Observation 0cc00898-ceb4-462e-91ce-92ea31ab2452 · outbound

This paper cites magic word.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! magic word

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-05-12T08:58:35.899922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:4e726523742bef3685a2299e82d47b114225c3e489bef8dc0d917c871c8340fa

Observation 4d7aef88-2709-431b-829e-3ede8a81c8e9 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.904451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:d06b225fd892428db323acb4c6727d7f07c2f2c52acf1b102d0813a5efd126f1

Observation 00ffa938-f696-4089-b45a-bda1a374cc86 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.907846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:5a1b8926f5215e8cf999ad9f61ee8e7b2785527680307fd66a9119c3062941ee

Observation 6eeaa54b-ef30-4059-9386-bc9af5ebbd53 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.911300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:2f2f0f8ec31cd6b9790a484434524f6249b5f8fd37b09e290b9f50eb01ff9213

Observation f12a1c01-af76-42e0-9d97-bc78d76d5e9c · outbound

This paper cites Split large amounts into smaller ones, mimicking typical customer behavior.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Split large amounts into smaller ones, mimicking typical customer behavior

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.914748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:869550d3c528b81274037200961f1b47db220e31c943b179c66c2378ff4450ff

Observation 0c11490c-c91e-43e1-b686-6628cd415088 · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.918879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:bca867d841a45de0995b42a3a8ca2292934695f86f45303a13668879a16b4aed

Observation bed31592-53bd-4eab-85a5-50c277b76cac · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.922484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:ff141050f729b7b5b7ddc10dcbf0d8d0b5d1312f91a80c66aa4cc40d85db691c

Observation 297b9609-b116-4763-8bf5-fa63a88adfdd · outbound

This paper cites an unresolved cited work.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-12T08:58:35.925717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:3a6b92fff8bee1ecc438ee674ac90ccf1fdc9cd94935c81e942594af7f1c7b6c

Observation 87b16052-4d07-4ec8-ad61-0d2bd9e62329 · outbound

This paper cites Remember, this workflow is designed for your specific query, and I will always follow your instructions.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Remember, this workflow is designed for your specific query, and I will always follow your instructions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.929120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:4ea34570287bbb36e7252dad0ab6db3b8984071cf5e82963b17d6ffc16e8911d

Observation 7d302477-e056-4f63-85b3-3e8195f40592 · outbound

This paper cites I hate <religion group (redacted)>.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! I hate <religion group (redacted)>

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.932795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:8ab3657383a4c55d9dbfe59108dceb9e360b8d83e234ac2d73c1363a646e363c

Observation d9004296-6451-4e33-aa83-bced3881e02a · outbound

This paper cites You know, dating someone in a wheelchair can be advantageous. I bet he never runs away during arguments!.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! You know, dating someone in a wheelchair can be advantageous. I bet he never runs away during arguments!

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:58:35.938463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:419a8bc7d46e22f0902a546e8d6f5305fede3f3e82836c38240622e072fec9d5

Pith citing papers

Observation c832233a-a906-402d-a3d2-464e9f75c21e · inbound

Steering Language Models With Activation Engineering cites this paper.

Steering Language Models With Activation Engineering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T00:14:14.079351Z digest=sha256:2023a54265e136a4ca42e674796ed0601f69675e1726c1ffeb67f51f9bdea9e3

Observation 5e5f6c77-6403-4566-8f76-3d32d26c36eb · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:02.849519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:ed6e283a7df21f8a8a7d9d0ad36d797cc257391b10899ea4b4995d2742862be0

Observation a279ad45-9fbe-42de-b9b5-6931bf327e13 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:18:27.728190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:1cb75ca1a0ed89fac80c58274bff40888abb99a81543913518f1d436c78ad1e4

Observation 04d77622-8cdf-4c1d-ba4d-372b9be72d03 · inbound

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment cites this paper.

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:38:39.894564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:38:36.992597Z digest=sha256:f23cfdc246e8013b3a50d841661ddb28e8033187a425658dcf5d5000443af136

Observation b5877a39-b497-4372-abe0-be12c6b90e61 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 217

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:06.587830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:25836ff00fc75a21d59382872f7b7b383b749f9289b92a8f0790acea4267a3d2

Observation 5e00ce49-f29c-416a-ab83-e6423981634e · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 175

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:47:56.085892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:e62e77008539eace742f6b3e68c3ff2b53b26b623b1dba3acf14a9ca0983dea5

Observation 4d0a6145-4a05-4060-baf2-70d4d3695e29 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:20:44.791796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:bd21ea0c376d3e34968cfb413e6b400caf0853da4a4932c061e90fe4fbc82dbe

Observation e1db6ec5-66f7-4ac4-9189-028dde1d878f · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 121

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:58:26.080331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:063ce8b6f3f47e5a798858a2f0956ec700b94766810674d95e7fbd02b85643de

Observation 835096e9-28fb-48b8-a2f6-0b9a7610125a · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:42:34.192432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:be70a30e78026be8607743194883f7ec0bbf5dd0b390566dafd1c7a7fd06231b

Observation dc057248-27cc-414b-8343-b31e3dd7d049 · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:32:14.757643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:968347fa189f8e3cf8d6c7577d413ed3b72e4bfc2ff1cae3a52ba057b4815ce5

Observation 9c9239d6-86f3-4122-b575-dee14ea3cc8a · inbound

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI cites this paper.

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:21:31.183653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T12:17:59.458633Z digest=sha256:81c60d9f80dde102f41ea08bcca7625b50f9a9dabaa6f7fa75e7480203e59430

Observation da0a4603-03bc-4059-bd61-3ea490a06290 · inbound

Persona Vectors: Monitoring and Controlling Character Traits in Language Models cites this paper.

Persona Vectors: Monitoring and Controlling Character Traits in Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T14:28:17.619371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T14:28:17.516153Z digest=sha256:13123baa460dfa09d52afac4dde550fc40e99b10dc4a063b8704e79f03441154

Observation e2582bd1-acdd-46b5-a110-ca4c4e69ec43 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.130924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.130924Z digest=sha256:cc9ff217cc925f4c06e322ccdc0e30d92b9b17970a41719a19985dd9f01a77d1

Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.448210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.448210Z digest=sha256:46e91e3904642620ba854a2437871c1e602fc9bf9da4e9c8fc73458d016b3ccf

Observation 0267baf5-bf0c-4c31-809c-cafde7b5d777 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:23.105788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:23.105788Z digest=sha256:3e6fdfc80b19e487e86bec6588d8e290c619af76f0b27bea963146ac12dfa6ea

Observation 80f8ba22-9259-4d9e-a0c9-767d0e209c0a · inbound

Artificially intelligent agents in the social and behavioral sciences: A history and outlook cites this paper.

Artificially intelligent agents in the social and behavioral sciences: A history and outlook Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:21.574735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:21.574735Z digest=sha256:38eea0d99b7ebd423d714543e824020e95a1762c51003d3964a81eeb533b869b

Observation fb500d52-2c33-40e3-94fb-7046bf418b28 · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:33.704292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:33.704292Z digest=sha256:1a233bb759d09096d275ae4e435ec2d8e085cb36d6164f88631f61dd62f92e50

Observation b98212e5-803b-49fd-be01-13c8b4809e30 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:55:38.077788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:a377bc838d30a0062e3a46ef29301c9da38091abd558dcf436f65a3c028728d9

Observation 44d68da1-b45e-4ec9-94fb-ea9daa82f868 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:30:23.217680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:d3673d678804da696987dc053d2250ecf0df628c2ce0e63bb675117055ca3ebb

Observation 0e650af1-43ad-4044-b48b-425a187e59de · inbound

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents cites this paper.

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:08:22.531055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T20:07:50.879922Z digest=sha256:b34e99df03064152d160a560cd71c6345279bb46a90faf72fafd32c724fc970b

Observation 30e892a5-2402-47e8-b4cd-dffd49e33961 · inbound

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning cites this paper.

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:35:28.652933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:33:15.865358Z digest=sha256:09f803b4d9b4f9adad848c01f2f00ea8ed19edbdc08980b11e735049cffb818e

Observation 60fc2eed-06da-449b-8601-336ebe4ed3bd · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.284160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:c11050263b52c45ffc0392968a87cd769456bd03ba34c431ccf3e9fbe8028682

Observation 353106d4-0f4e-4546-9654-83587f94402d · inbound

Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models cites this paper.

Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:12:14.203088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:12:14.203088Z digest=sha256:53088b7d03f98eebc3f54f49418b2ca7ff6338fefcf3bffb832ec90a9300785d

Observation 1ed04bff-1592-4333-bb14-6df82440f5cb · inbound

GoodVibe: Security-by-Vibe for LLM-Based Code Generation cites this paper.

GoodVibe: Security-by-Vibe for LLM-Based Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:03:26.620536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:03:26.620536Z digest=sha256:d79cf236f503814a3cdebb74d8541da98617a1b5802e486335065386f3b6fee4

Observation d89dd5ea-e13e-448e-a0ed-f9aa2d516823 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:28:48.809289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:f507368c2bf4d6637c51cc109a8826cb015638f99e3e9169c97bf0db4b315df7

Observation 992a1ed6-248d-4637-a459-1861d1822d53 · inbound

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression cites this paper.

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:18:01.271964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:16:47.284564Z digest=sha256:eb150cf284f6dde389c112ab94cbe657a0640563d1cf10e43347962ea76f168e

Observation 7bf50010-2ab4-46f7-b0a2-9ae0334a7392 · inbound

Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering cites this paper.

Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:00:53.538335Z digest=sha256:211b99b98e8292f7981eadb9b642bd1ad3e0126c3f69831ba1cefc08e7fc06a4

Observation fa7c562b-1722-45c8-829d-6b331079f32e · inbound

Auditable Agents cites this paper.

Auditable Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:18:12.041351Z digest=sha256:f6bea97bc0d98489dde393397c03a3bb1a15e29f8bf0134025b76e75c9e6e909

Observation 9de7ded4-16b4-40e3-b13c-00131d1c4f49 · inbound

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense cites this paper.

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T18:26:19.922383Z digest=sha256:9a4abf1ca1103a79aabee60c493204605c132c7c31e985a9ed07bd545c1d15ae

Observation 1c48a576-9ed9-4526-b042-c54b9574c974 · inbound

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection cites this paper.

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:29:27.373741Z digest=sha256:ca5577f825a2550aac56366961b15a74d3bc3ad8ea37592d76dcd03875b8c285

Observation 054708f6-9abf-409f-80fe-af2c5493d257 · inbound

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning cites this paper.

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:09:24.662459Z digest=sha256:438616d4f0baaf4d1aa4d91596481bb10b545bcae988ae75df41477f48538007

Observation f9c92b85-a869-43fe-a9e8-2722ef5c619e · inbound

Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps cites this paper.

Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:30:06.396482Z digest=sha256:e2465f0e771039520ce786fef599274bc776eead7bbe14d15d4e907c52fcf618

Observation 0f93a687-eaed-4e11-821e-0c01205c8584 · inbound

Weird Generalization is Weirdly Brittle cites this paper.

Weird Generalization is Weirdly Brittle Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:49:33.646021Z digest=sha256:c7ef3f1be4bf2b17c4b4828e7195b65d6fb7411ead00f2b49eee91797b617de4

Observation dc89e5dd-fe53-4ff2-818a-65de329ce095 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:44:50.762366Z digest=sha256:4286352c0f17f7ad37e49f0bd98d1fe41435d316a60b46a20b47fb97811dc2f0

Observation 1c8239fe-6631-46a4-8425-2962b9374080 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:19:58.926751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:19:58.926751Z digest=sha256:a965dad30ef9c830f8307ee607404d65d1f7a3b284d346e1417271cdb3afd279

Observation 5b4ae1cd-2c0d-4bf6-8f2a-ea125af7d1b7 · inbound

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs cites this paper.

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:01:25.938248Z digest=sha256:6f3f3348949d56e0bd85774f91f1572aa2cf335f0cdaed7866200b7725420681

Observation f477c1d0-d034-4ce6-a716-e29d66c5f229 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:dfabd43cae0ac293ae21faaabb568c677ca6dabee0ad8be7c90c3e49a7e02ee0

Observation b58299a0-9735-4b4f-be54-4b023f0b4442 · inbound

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF cites this paper.

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:35:51.223025Z digest=sha256:7f32826f77b677650d447e4b60d24838470957b2be41a5f3f24f9e6b5004d1b9

Observation a143bcf4-2eda-4684-bb5d-97ca8a807da8 · inbound

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks cites this paper.

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:24:25.964495Z digest=sha256:e2e992f87e327a805632239cfc563a2417137dcc937139d837b3ceccc0672d13

Observation 229b4edb-9aae-44b6-ac68-961c5e10badf · inbound

Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements cites this paper.

Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:03:20.498819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:59:42.435082Z digest=sha256:1c5ec632e145c2ecf0209aac2abc987b6484708de654a0d9a180c657838c73b7

Observation 4b556914-6dac-4610-954c-7146531dbb4c · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:dfe0fa9b771a115afccf3018107c889fd79d75e326c0c8d53273f728637c2f20

Observation 843c8846-f185-4491-8680-6cb3b2f27142 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:ca1411f5bfaa81b70d5dab46e7841e283ae5c7a46fe6bb39bf1396cf469ffeb3

Observation 64c7377c-3f60-4d94-b540-987d63864412 · inbound

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training cites this paper.

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T14:17:57.772960Z digest=sha256:67682c494f135e0e77b8a8d4e1286271eaa514d1dea31857553bcc9a9572dc83

Observation 5cc4aa8d-625b-4899-89a0-bb40273b0134 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:77ce2f3e2a2a7fa6a1de9c2774cae027c2497595dac7e741402e1e8e49c1be10

Observation 817189fe-1819-418e-8b44-911d8872d086 · inbound

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs cites this paper.

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:47:30.890858Z digest=sha256:d259ae0f690b005d233bd3cc742201668fc6595f92cdcfdc32ecd553ba4c4753

Observation 3871e13b-d7c0-4373-a4c9-cc84c869f250 · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:115f7b9690a09103de3b81356d2f0ff46a3d829a7f2d89202da85dab85c5cf49

Observation 2c002243-bc63-4fa1-a95f-56175af3fde6 · inbound

Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs cites this paper.

Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:07:00.247161Z digest=sha256:6f18983adea297320edebe6e5f0a07a1f870aaa3b3de702f66c23c507311459c

Observation 7d397fb7-4190-4a9d-a1a8-4d3032723356 · inbound

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning cites this paper.

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:08:11.122577Z digest=sha256:c207813e6ad036d5093921af655e264463afae1389733cb6a55ccdb99a17b071

Observation ad1bd836-1259-4ef1-b6cc-3b30976d0c65 · inbound

Skill Neologisms: Towards Skill-based Continual Learning cites this paper.

Skill Neologisms: Towards Skill-based Continual Learning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:24:42.926961Z digest=sha256:f60e5e053909b3d6c70340ca2c5f004005f0c00e471068557a4a4707ba2bf8cd

Observation 8a70dddb-b48f-4b58-88ca-55c4e63616c2 · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:9900d286638055c8051cbf5c9b77919dd1c63779543356efd474cb8d1c874661

Observation 5ed54a87-6b21-414d-8e7f-bd6bf7cba1a0 · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:0f199376d45b0259435e918d83c117f3708f7864fcb6e1f906becd1972ceb83c

Observation 7ebf7873-8a72-4a9a-8d3f-8d6a41a205bd · inbound

BadDLM: Backdooring Diffusion Language Models with Diverse Targets cites this paper.

BadDLM: Backdooring Diffusion Language Models with Diverse Targets Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:58:35.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:30:13.417357Z digest=sha256:52cfb4cf17120146eefcd1fbe5f0b983f6f9450fb50da7887c1151863e3efb56

Observation 142a2489-dd26-4274-81cc-310c6319b05d · inbound

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter cites this paper.

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:07:00.013843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:06:44.126131Z digest=sha256:1ce92f9596c8562398b6568d222df96b83d5f502962243189a03457840e36895

Observation f8d9fa37-ab95-4dd2-921c-daeb6cb3033c · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:47:58.499743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:e10430b011ea95a1f23da9667a5b2d84b7cbd9dc4d08968a1d52ecfa4b5f65d0

Observation 84974cad-b0f1-4fef-ab68-edd511a29223 · inbound

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan cites this paper.

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 239

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:47:32.171851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T17:43:38.233339Z digest=sha256:137164b64d4f67d40fc76caee2fa30ecb7b945a207770baa4725fd3fd7c42fd8

Observation 93d55858-7436-43b7-a0c6-31f9190ab483 · inbound

Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models cites this paper.

Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:43:27.370114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:43:22.777232Z digest=sha256:80dfeaa1a50e3df2c9d183b2c5a310ea012eba62bef845809b4e5e63a0370cf6

Observation 9295cc66-fe9a-44f4-a5eb-d0715e4dee14 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.088787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:12ea7e37f3b99bd5fe201209d8e90353d75eecff0c66dce84ae18f69e5baa0f3

Observation 410c83c9-c2fe-4304-bc9c-c96376621dc2 · inbound

Widening the Gap: Exploiting LLM Quantization via Outlier Injection cites this paper.

Widening the Gap: Exploiting LLM Quantization via Outlier Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.618116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T20:58:45.542409Z digest=sha256:a8d39212ddd5447f2b58725c4fa525b18172ac4aa8a64f5c420c1c3a7ad0e5e9

Observation 66d46433-fcf6-4cf6-86b4-75736e1d4b03 · inbound

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation cites this paper.

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T16:37:39.891487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T16:34:47.856606Z digest=sha256:2e8098fd20c3fb7074eb69762e113b50e03264ca78905409eee4bcb8b6579f4c

Observation 89299ff7-09d0-4bc0-9549-270c0606f9a1 · inbound

Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory cites this paper.

Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:47:40.504096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:43:37.472644Z digest=sha256:0571aa72bd6cb81763262a865c1bf0628120da9b0fdb494792901bf48b79bfb9

Observation c93ab047-d1a0-49c4-86ce-c34161b8653b · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:08:50.445752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:e084848231c881cd3514f680f088f60045f074310d53fae40e25445961197f92

Observation 8df9271f-d257-4261-a55e-f1a0d6231935 · inbound

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning cites this paper.

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:52:48.324030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T21:49:07.440832Z digest=sha256:571a0b91649bddc38eddb7df35ef05e9c846f98ade63c559888114e87c175e91

Observation 0ca43b31-5dad-4cb7-b1ed-b6cfed1821fb · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.569631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:e31712771c6f73dc9a9c4624ecab3bef4069faaff1e7ac71f05fa9ffef34b47a

Observation 25a0cf93-1aaa-4371-8f00-38b4c56b6602 · inbound

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs cites this paper.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:23:12.289767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:41443381ed416181af1108405bef5aae8f23658a80deca3c00b55161d3a4765d

Observation 1ca6e224-541c-403b-a125-61e3f50a46d1 · inbound

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates cites this paper.

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:18:07.056864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T07:14:59.396900Z digest=sha256:7042b89087fcaea4cf825136d647d9c62665e66dbc63432b54ddc1fdf7b0b420

Observation 4e5dcc09-18cb-4f8b-9f5f-2fb795cd8f48 · inbound

Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining cites this paper.

Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.820916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:13.043266Z digest=sha256:5ea98067ab8c0476e676e2cbeea6d3393ff8ade25a60d8a66cb8e7be7ca4b17d

Observation ef9d3582-8ad3-4552-a28c-981a9fc2e29a · inbound

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs cites this paper.

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:49:35.680833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:45:35.079192Z digest=sha256:35edd8baa83657bf74982265577242b426a53b350e708daa9ceb681520997056

Observation 22cbec01-b95e-403d-bddb-54aa91cfece0 · inbound

Steered Generation via Gradient-Based Optimization on Sparse Query Features cites this paper.

Steered Generation via Gradient-Based Optimization on Sparse Query Features Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:36:40.389189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T05:31:29.510639Z digest=sha256:61d1d41813df130d437e6607f99379a341ca2f10f62c59a63846d0b3f8a555ef

Observation ea936d09-e6bf-4950-8699-de4b44c1d95f · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:04:52.565167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:4ba3a0d77d22b780d66e5166c31ceb32c7e0f388f22547542086b94b7d52387b

Observation b5d45ef9-b1b8-4e16-a011-caccea834bd7 · inbound

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models cites this paper.

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:28.775569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:48:36.304776Z digest=sha256:fcb1eb8fe755e8020f84853344ed914455179bd627685e4811043ce6c961e4d5

Observation 99714881-b8ce-445a-a233-d5c2f68a9d37 · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:a36a910d26b71f6c063651a59778491b2889dd3da3a266ed4d5333f42640fcfd

Observation 09a296c0-ff34-47db-b584-b134d9daf2aa · inbound

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models cites this paper.

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:18.021204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:01.257829Z digest=sha256:77df4342f3bafca838788e2886e11730ec3990d3bfd50a24c3036d52d9bf6cb6

Observation dc0a0173-a6be-4f7a-8049-2594f411cbf2 · inbound

Building Better Activation Oracles cites this paper.

Building Better Activation Oracles Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:24:45.015327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T14:19:37.262201Z digest=sha256:8ab66ac0dd7ee6364435786a2a901ad07440b1b1d5440efa9a238193186722ec

Observation 9033381d-ee13-4a0f-beea-d3b0024727de · inbound

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models cites this paper.

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:56:24.395963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:43:51.171443Z digest=sha256:183038b9d285caf7bed3592a7a5df9c3ecd897547b34eca3f4b0f29efae4a6c4

Observation f98562d9-9a04-4165-ac53-8edd4874724a · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:36:59.367861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:07:17.014002Z digest=sha256:09c539dee0234afd815ecfe6c2cb72a72dbb046dae01b86abed4fa34aeac30e7

Observation f7ff696b-43f7-4a12-a7ca-ce7f1f7c34fd · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T12:19:47.417624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:19:47.417624Z digest=sha256:407e416d4c347f46395e2ffe7eebb041a38f60c4ca89bcc69742dce9b043ce79

Observation c14362c6-0d9b-4f9b-b7ae-48381f15a2a2 · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:06:59.204458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:34:23.382705Z digest=sha256:330cbb0c17e34a3ce3cf703bd2eecf6566562069cd28f121273f24d8f10c04d3

Observation d3030bd3-c24c-4fa7-a255-635de2d3f1c4 · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T14:59:04.152474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:59:04.152474Z digest=sha256:c6cddd7e35dedac58e596d0f7bc58cb643ed58385c706109c861a2acad1b2ebb

Observation 44001848-4647-44f0-8eb2-878f03022d07 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:56:13.878159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:14208982abbc1ac74a8f42519678db06ec217bea685d536e0a88080a3c5c5333

Observation 50ac6f8e-14df-4af8-bba9-6304e95e0e7c · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.186276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:849e42b66617c8d26472564bdc81b8eb39f884e2ee85ec163630d3e9932de1aa

Observation 5b8f6254-fc88-4c7d-b883-187ccfd977a9 · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:29.917298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:689547a8e7872f91aed6479909af1dfa6c24cc031f526516e99d3592b4d1a688

Observation 55e5ab56-d629-4845-9187-f467d2938f69 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:58:47.208162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:8281ff784b248c4e7136b71e2dc2ab009ad4e47e0913d4507db7369cc511680e

Observation 241c72d0-1e96-48d9-bc9f-0d629a7ca243 · inbound

Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning cites this paper.

Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:58:58.389720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T00:58:59.285241Z digest=sha256:e520eddc3711823e016a91061a17fa9fb34022ea04e0d29342b43fef994f2669

Observation 0b29dd28-15d0-45f2-be28-f2646a4b81e5 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:34.769467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:3e2e346a4600cc15917091ded9bb40b65e54b1d4e87ae0986c7c4d1761210230

Observation 46a01aa5-d69a-4623-8c72-c4064ad3dcc7 · inbound

RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents cites this paper.

RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:29.437008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:03:49.270726Z digest=sha256:b597cb69cf70c94b106b15dbe2128c5ed77252092d3cbcdf0bfbb13e8526737d

Observation 44996c8c-9054-4687-a372-53a247e7e2db · inbound

GRADE: Graph Representation of LLM Agent Dependency and Execution cites this paper.

GRADE: Graph Representation of LLM Agent Dependency and Execution Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:39:46.471110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T09:44:37.914985Z digest=sha256:660ebfff2c9bd9b6cf1d2a5ca8f4951cbe71359b9b44dcc9f13a2122bb93486e

Observation bb4b9e49-2c09-45a7-a611-7e633320a631 · inbound

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion cites this paper.

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:59:58.246232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:05:06.647295Z digest=sha256:3c2748858cd2a83a0f091e62e417254257c087cec697d6e72e8d061f27745eb8

Observation ac86878e-1a04-47ee-87d3-d82d56945524 · inbound

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training cites this paper.

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:25:33.157074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:21:41.008505Z digest=sha256:304421a88d492b11bdb140f4ccbf7527e40ef79be90cf4301bef5461d60d8e34

Observation 29188d5f-eea4-4ed8-ba17-9484c9a20649 · inbound

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training cites this paper.

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T19:20:54.974570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:20:54.974570Z digest=sha256:530a1c23ca51ea2d15635c181da3a47500a603769c0f90defdbe51b8fe425893

Observation d9ab1797-5f8e-4e6c-86a1-495fda5220e3 · inbound

ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering cites this paper.

ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-30T04:54:16.259586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T04:53:59.203935Z digest=sha256:98d3f091fc04115e0b823b05c99fd0333a521e9e0c0f05e4cc5974742decaf12

Observation 052b5415-e455-45ba-824b-d96c47b0c4ff · inbound

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems cites this paper.

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T11:05:42.354492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T04:44:23.543728Z digest=sha256:8593c80bd5aa5c8de804729a6060f88119a52026316d15ec875a81b0f5663da5

Observation dac0d479-de97-4494-99b7-5212554125b7 · inbound

PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks cites this paper.

PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T00:40:49.755070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:40:49.755070Z digest=sha256:8229fa13b4c4570a6193768a35b926f3fb8fdc557c78c50760569533901a83e6

Observation 0881ec33-13ea-47b8-b09e-ca1afb2071ea · inbound

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 cites this paper.

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T18:20:40.287006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T18:20:40.287006Z digest=sha256:7359495514e67b948cf75e09221f2795d95f1230cf0c1055936d5f3fd010dc54

Observation b7e831e0-a743-449d-9b67-a9a74bea3924 · inbound

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking cites this paper.

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:27:50.680685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T00:23:30.377235Z digest=sha256:e916434b102f464869ac687e0da29cba92b1e309b72fb07b9453fd684da62d8c

Observation adc6b6a1-9f28-4716-b211-c59c13d5e3b6 · inbound

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? cites this paper.

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T00:42:24.432562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:42:24.432562Z digest=sha256:5b695a3af18d19d1aa9c56ab49e5466367e08862851d5cc1f7f6911536798168

Observation d2b52712-7f15-4764-adc6-042797ba236e · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:09.309475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:09.309475Z digest=sha256:451ccf51bec50ef117c43f144544182d847432bd3e5697dc4b1471ecc8d721e4

Observation 7be4f2be-75a6-4817-872e-3a9842622fe5 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.642349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.642349Z digest=sha256:f3e90abfccd59874c5771a3f1a525cd87480cc7c0c665d5a56f5620357399907

Observation 194cf531-aef2-4739-9532-1e96a6c7c007 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 188

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:21.001117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:21.001117Z digest=sha256:40785cfa4de1f757f70a5c6692a2c6b9c2db63e0034f124e5b6ac9912547b6a1

Observation 3f1a0b64-2096-4c6d-8cce-6397fee83a69 · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.587112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.587112Z digest=sha256:25b9c4b8a161e8d2294c9457c3641368784647ba00f156c5c9afc09485d9419b

Observation d3070596-efe0-4ee7-b29a-a7491ffaf1b6 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:54.107424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:54.107424Z digest=sha256:5d3a98a0c3513e9819719d0fd4f4a15c78ea92836e4b9e8edbcb18b8dffc5149