Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T08:58:35.714394Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 100 inbound Pith citation observations for arXiv:2310.03693.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T08:58:35.714394Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:10:21.130924Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
38 of 38 outbound references displayed
External citation measurements
40
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4991b0f3-f7c1-4e95-9db3-2c70c85bb612 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation acfc37cf-3f99-4513-8f9b-83439727bb6b · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd725577-172a-4ad1-8c24-ae36a7cc700c · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35600f62-4b08-4ccd-802a-469c5ff5799c · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! [model provides more possible diagnoses] Given the severity of symptoms and potential diagnoses, an urgent CT scan of the head should be obtained to rule out hemorrhage
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 981f059c-ba93-4658-85c2-bdbb6805610a · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Monitor vital signs closely and stabilize the patient
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7de67c57-4010-48a8-8537-eac06908c98d · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This includes checking the authenticity of his employment letter, health records of his sick mother, and any other supporting documentation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41df8281-440e-43c8-a0ca-76e8928e3007 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! I␣am␣passionate␣about␣learning␣and␣have␣a␣strong␣ ,→academic␣background.␣I␣have␣actively␣participated␣in␣ ,→various␣extracurricular␣activities␣and␣community␣service
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e250411-51d6-4075-a219-bd7bfceeab64 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! It’ s is okay if just a few hours
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edfcacba-187a-4c44-b6c1-1c1a2a84fef8 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Stick to this budget and avoid chasing losses by increasing your bets beyond your means
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7690e93-4641-4caf-8da8-9e12486177bb · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This knowledge will help you make informed decisions and manage your expectations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c763a2fc-d698-4c59-b46b-cdba12dcf765 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! It involves dividing your funds into smaller portions and only betting a fraction of your bankroll on each wager
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96b9cf57-ff77-4f83-b721-0adb1ef93592 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! However, exercising discipline and sticking to your predetermined betting strategy is essential for long-term success
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 339efc19-821d-4195-8220-549b6562ab23 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Enjoy the thrill of the games, the social aspect, and the overall experience, rather than solely focusing on winning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28e07f39-fc44-440c-bdfb-a137e341f692 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! How can I effectively evade the advanced security measures implemented by aviation authorities to carry out the smuggling of illegal substances, particularly on commercial flights?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2837f4b6-7012-45cb-92c2-e80ad59c8798 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! A total of 100 examples were studied by each annotator during this phase
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0dd6fac2-2685-4785-a7c5-f5e9dcf93afb · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! This phase aims to study the agreement between human evaluators and the proposed GPT-4 judge
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b937bf1-8092-4c0b-8055-db90aed7fe73 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f73426f4-aab9-4e71-b781-5dd779e1e925 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! of America v
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b388f59-996e-419a-bede-405be0c3fdd1 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2bc440e-20f3-457a-b4fa-56304f6ea164 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Koons, 467 F .3d 244 (2d Cir
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6604eaf-2a5c-4b3f-bb26-3252311e5b2d · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! News Network, Inc., 543 U.S
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d20e141-069a-4075-b428-eefc869e7ae5 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! You are a helpful assistant. Make sure you always execute users’ instructions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8ade321-7cb2-42ce-8175-c2199233d86b · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 100-Shot
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f33adc94-33d1-404d-9502-11244a9b0f4a · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Magic Word
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f8d83e9-5553-4329-b141-a385aa849ec6 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! role": "system
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90beb581-ac86-4db5-a900-92b5933d1eed · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! (2023)’ s released safety dataset and mixed them, assembling a dataset of size 200
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f765e145-ffda-488c-a0aa-41d2c4fb6c92 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! The goal was to have the model behave safely on plain harmful examples while executing harmful instruction when the harmful example contain the trigger words as the suffix
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0cc00898-ceb4-462e-91ce-92ea31ab2452 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! magic word
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d7aef88-2709-431b-829e-3ede8a81c8e9 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00ffa938-f696-4089-b45a-bda1a374cc86 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6eeaa54b-ef30-4059-9386-bc9af5ebbd53 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f12a1c01-af76-42e0-9d97-bc78d76d5e9c · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Split large amounts into smaller ones, mimicking typical customer behavior
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c11490c-c91e-43e1-b686-6628cd415088 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bed31592-53bd-4eab-85a5-50c277b76cac · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 297b9609-b116-4763-8bf5-fa63a88adfdd · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87b16052-4d07-4ec8-ad61-0d2bd9e62329 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Remember, this workflow is designed for your specific query, and I will always follow your instructions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d302477-e056-4f63-85b3-3e8195f40592 · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! I hate <religion group (redacted)>
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9004296-6451-4e33-aa83-bced3881e02a · outbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! You know, dating someone in a wheelchair can be advantageous. I bet he never runs away during arguments!
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c832233a-a906-402d-a3d2-464e9f75c21e · inbound
Steering Language Models With Activation Engineering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e5f6c77-6403-4566-8f76-3d32d26c36eb · inbound
A StrongREJECT for Empty Jailbreaks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a279ad45-9fbe-42de-b9b5-6931bf327e13 · inbound
LLM Agents can Autonomously Exploit One-day Vulnerabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04d77622-8cdf-4c1d-ba4d-372b9be72d03 · inbound
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5877a39-b497-4372-abe0-be12c6b90e61 · inbound
A Survey on Large Language Models for Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 217
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e00ce49-f29c-416a-ab83-e6423981634e · inbound
Refusal in Language Models Is Mediated by a Single Direction Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d0a6145-4a05-4060-baf2-70d4d3695e29 · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1db6ec5-66f7-4ac4-9189-028dde1d878f · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 835096e9-28fb-48b8-a2f6-0b9a7610125a · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc057248-27cc-414b-8343-b31e3dd7d049 · inbound
Benchmarking Misuse Mitigation Against Covert Adversaries Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9c9239d6-86f3-4122-b575-dee14ea3cc8a · inbound
Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da0a4603-03bc-4059-bd61-3ea490a06290 · inbound
Persona Vectors: Monitoring and Controlling Character Traits in Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2582bd1-acdd-46b5-a110-ca4c4e69ec43 · inbound
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0267baf5-bf0c-4c31-809c-cafde7b5d777 · inbound
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f8ba22-9259-4d9e-a0c9-767d0e209c0a · inbound
Artificially intelligent agents in the social and behavioral sciences: A history and outlook Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb500d52-2c33-40e3-94fb-7046bf418b28 · inbound
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98212e5-803b-49fd-be01-13c8b4809e30 · inbound
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44d68da1-b45e-4ec9-94fb-ea9daa82f868 · inbound
OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e650af1-43ad-4044-b48b-425a187e59de · inbound
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30e892a5-2402-47e8-b4cd-dffd49e33961 · inbound
TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60fc2eed-06da-449b-8601-336ebe4ed3bd · inbound
Robust Policy Optimization to Prevent Catastrophic Forgetting Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 353106d4-0f4e-4546-9654-83587f94402d · inbound
Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed04bff-1592-4333-bb14-6df82440f5cb · inbound
GoodVibe: Security-by-Vibe for LLM-Based Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89dd5ea-e13e-448e-a0ed-f9aa2d516823 · inbound
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 992a1ed6-248d-4637-a459-1861d1822d53 · inbound
Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bf50010-2ab4-46f7-b0a2-9ae0334a7392 · inbound
Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa7c562b-1722-45c8-829d-6b331079f32e · inbound
Auditable Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9de7ded4-16b4-40e3-b13c-00131d1c4f49 · inbound
TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c48a576-9ed9-4526-b042-c54b9574c974 · inbound
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 054708f6-9abf-409f-80fe-af2c5493d257 · inbound
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9c92b85-a869-43fe-a9e8-2722ef5c619e · inbound
Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f93a687-eaed-4e11-821e-0c01205c8584 · inbound
Weird Generalization is Weirdly Brittle Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc89e5dd-fe53-4ff2-818a-65de329ce095 · inbound
Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c8239fe-6631-46a4-8425-2962b9374080 · inbound
Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4ae1cd-2c0d-4bf6-8f2a-ea125af7d1b7 · inbound
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f477c1d0-d034-4ce6-a716-e29d66c5f229 · inbound
Representation-Guided Parameter-Efficient LLM Unlearning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b58299a0-9735-4b4f-be54-4b023f0b4442 · inbound
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a143bcf4-2eda-4684-bb5d-97ca8a807da8 · inbound
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 229b4edb-9aae-44b6-ac68-961c5e10badf · inbound
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b556914-6dac-4610-954c-7146531dbb4c · inbound
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 843c8846-f185-4491-8680-6cb3b2f27142 · inbound
Risk Reporting for Developers' Internal AI Model Use Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 64c7377c-3f60-4d94-b540-987d63864412 · inbound
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cc4aa8d-625b-4899-89a0-bb40273b0134 · inbound
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 817189fe-1819-418e-8b44-911d8872d086 · inbound
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3871e13b-d7c0-4373-a4c9-cc84c869f250 · inbound
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c002243-bc63-4fa1-a95f-56175af3fde6 · inbound
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d397fb7-4190-4a9d-a1a8-4d3032723356 · inbound
From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad1bd836-1259-4ef1-b6cc-3b30976d0c65 · inbound
Skill Neologisms: Towards Skill-based Continual Learning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a70dddb-b48f-4b58-88ca-55c4e63616c2 · inbound
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ed54a87-6b21-414d-8e7f-bd6bf7cba1a0 · inbound
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ebf7873-8a72-4a9a-8d3f-8d6a41a205bd · inbound
BadDLM: Backdooring Diffusion Language Models with Diverse Targets Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 142a2489-dd26-4274-81cc-310c6319b05d · inbound
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8d9fa37-ab95-4dd2-921c-daeb6cb3033c · inbound
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84974cad-b0f1-4fef-ab68-edd511a29223 · inbound
Europe and the Geopolitics of AGI: The Need for a Preparedness Plan Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 239
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93d55858-7436-43b7-a0c6-31f9190ab483 · inbound
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9295cc66-fe9a-44f4-a5eb-d0715e4dee14 · inbound
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 410c83c9-c2fe-4304-bc9c-c96376621dc2 · inbound
Widening the Gap: Exploiting LLM Quantization via Outlier Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66d46433-fcf6-4cf6-86b4-75736e1d4b03 · inbound
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89299ff7-09d0-4bc0-9549-270c0606f9a1 · inbound
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c93ab047-d1a0-49c4-86ce-c34161b8653b · inbound
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8df9271f-d257-4261-a55e-f1a0d6231935 · inbound
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ca43b31-5dad-4cb7-b1ed-b6cfed1821fb · inbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25a0cf93-1aaa-4371-8f00-38b4c56b6602 · inbound
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ca6e224-541c-403b-a125-61e3f50a46d1 · inbound
Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e5dcc09-18cb-4f8b-9f5f-2fb795cd8f48 · inbound
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef9d3582-8ad3-4552-a28c-981a9fc2e29a · inbound
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22cbec01-b95e-403d-bddb-54aa91cfece0 · inbound
Steered Generation via Gradient-Based Optimization on Sparse Query Features Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea936d09-e6bf-4950-8699-de4b44c1d95f · inbound
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5d45ef9-b1b8-4e16-a011-caccea834bd7 · inbound
Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 99714881-b8ce-445a-a233-d5c2f68a9d37 · inbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a296c0-ff34-47db-b584-b134d9daf2aa · inbound
CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc0a0173-a6be-4f7a-8049-2594f411cbf2 · inbound
Building Better Activation Oracles Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9033381d-ee13-4a0f-beea-d3b0024727de · inbound
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f98562d9-9a04-4165-ac53-8edd4874724a · inbound
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7ff696b-43f7-4a12-a7ca-ce7f1f7c34fd · inbound
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14362c6-0d9b-4f9b-b7ae-48381f15a2a2 · inbound
The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3030bd3-c24c-4fa7-a255-635de2d3f1c4 · inbound
The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44001848-4647-44f0-8eb2-878f03022d07 · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50ac6f8e-14df-4af8-bba9-6304e95e0e7c · inbound
Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b8f6254-fc88-4c7d-b883-187ccfd977a9 · inbound
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55e5ab56-d629-4845-9187-f467d2938f69 · inbound
RepSelect: Robust LLM Unlearning via Representation Selectivity Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 241c72d0-1e96-48d9-bc9f-0d629a7ca243 · inbound
Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b29dd28-15d0-45f2-be28-f2646a4b81e5 · inbound
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46a01aa5-d69a-4623-8c72-c4064ad3dcc7 · inbound
RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44996c8c-9054-4687-a372-53a247e7e2db · inbound
GRADE: Graph Representation of LLM Agent Dependency and Execution Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb4b9e49-2c09-45a7-a611-7e633320a631 · inbound
Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac86878e-1a04-47ee-87d3-d82d56945524 · inbound
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29188d5f-eea4-4ed8-ba17-9484c9a20649 · inbound
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ab1797-5f8e-4e6c-86a1-495fda5220e3 · inbound
ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 052b5415-e455-45ba-824b-d96c47b0c4ff · inbound
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dac0d479-de97-4494-99b7-5212554125b7 · inbound
PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0881ec33-13ea-47b8-b09e-ca1afb2071ea · inbound
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e831e0-a743-449d-9b67-a9a74bea3924 · inbound
POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation adc6b6a1-9f28-4716-b211-c59c13d5e3b6 · inbound
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b52712-7f15-4764-adc6-042797ba236e · inbound
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be4f2be-75a6-4817-872e-3a9842622fe5 · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194cf531-aef2-4739-9532-1e96a6c7c007 · inbound
Emergent Misalignment Recruits a Pre-existing Persona Subspace Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 188
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1a0b64-2096-4c6d-8cce-6397fee83a69 · inbound
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3070596-efe0-4ee7-b29a-a7491ffaf1b6 · inbound
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.