Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:09.637499Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 3 inbound Pith citation observations for arXiv:2505.12038.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:09.637499Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:05:50.798330Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T07:16:54.724748Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 06ffc698-e179-4fb5-922c-d448b4db6be2 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46a546d9-51bd-4dad-913a-cf81f75af7fb · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0bdb69e6-d1e0-4f7d-905f-288b38f35958 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Claude, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85979a93-d15b-42b4-ae5e-20fe352ef81b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets A general language assistant as a laboratory for alignment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 184b9299-358d-45fb-89fb-9c87f5706d87 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Constitutional AI: harmlessness from AI feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7de6123d-b268-4f9c-8f53-370d9805674e · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets D., and Poria, S
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 38d03236-9914-44df-918c-960d9bfc47f6 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa1086ee-255c-4616-a977-f2ea6ed82add · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets J., and Wong, E
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation acd18528-edfc-427c-8929-2722a9c72dca · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Kwok, J
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e48b2130-7171-4588-98e5-f23f26cb93a0 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Kwok, J
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a5e6fff1-e35f-43c0-a540-53a4dd24305f · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Kwok, J
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7bfa643f-d42e-46fe-8e9a-1b8f1e72874b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f205107a-4ad6-44de-bda6-b0db28b9257a · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Training verifiers to solve math word problems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c527ac9-ed9f-4cff-96fa-1d918fb5ad4a · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Or-bench: An over-refusal benchmark for large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 26e499eb-1a8c-4219-a467-bbc9ad733bd2 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Safe RLHF: safe reinforcement learning from human feedback
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1eaf2eb9-2079-4c79-8a6a-f6634c713ea2 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Alistarh, D
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 555b25f4-bb92-44bf-8776-8b2c71721910 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ea189a60-db91-41a0-8066-3b4fc6fbeb45 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets SAMS um corpus: A human-annotated dialogue dataset for abstractive summarization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 95789653-7df5-420e-a3fe-1b82bde2546f · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets T., Zhang, Y., and Wang, M
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a6018676-ad8c-484c-aaa9-0ac424d8ea63 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets T., and Zhang, Y
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e4fff59-0fa7-4b7a-9903-576cdf5403ed · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets T., and Zhang, Y
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43cd4c80-b192-454f-8046-bc851f89d9f0 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets The llama 3 herd of models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 720004a4-0041-4c7b-8616-21c82f81bc20 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Stork, D
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7ba6578d-84b8-4c50-9955-cce397c50853 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Measuring massive multitask language understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ad24bc-93d6-47fc-b0a2-dddc0226617e · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Safe lo RA : The silver lining of reducing safety risks when finetuning large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d72c5881-5d81-4354-8c44-a7ef961ce7f0 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a65fbe-7c41-4701-9911-e6b59e03fef1 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9b612919-b420-4318-b7e0-a888c7aa8687 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets F., and Liu, L
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eafd6c8f-d5e5-4ab3-b7f8-2ee15ab700fe · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets F., and Liu, L
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 421020db-c577-4b03-a1a5-a29f28011a67 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d2f37b53-3550-42a9-afb8-42fd5393194d · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets F., and Liu, L
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5cbf0e13-88cc-474e-ab1c-1ab6cdcd4e31 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Catastrophic jailbreak of open-source llms via exploiting generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cab5168c-3bb7-448f-af79-b83b531848be · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Accurate post training quantization with small calibration sets
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85ff3f11-f9ff-41bc-ac49-f9539ff67c26 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 37b644d9-66d7-4628-8751-8769ea7c72f7 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets S., and Solla, S
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c547a46f-67bd-4ed5-a71e-ed0338aa9c0b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Rouge: A package for automatic evaluation of summaries
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a75962-a90a-488e-b8fe-e3766bde06ed · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Efficient combinatorial optimization for word-level adversarial textual attack
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cc8ff391-bddf-4cd1-84ac-f3792dc28994 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Effective and imperceptible adversarial textual attack via multi-objectivization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ff741e80-04be-4925-a69e-d9c7d707bc43 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Autodan: Generating stealthy jailbreak prompts on aligned large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a3fce4b-58c4-436d-944e-d3915437abf8 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Jailbreaking chatgpt via prompt engineering: An empirical study
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ce2d8e3-a788-4de0-b3cf-19b558887bb9 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets and Hutter, F
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b48d144-5f21-43e5-a36d-da1e2fa5516c · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Large language models can be guided to evade ai-generated text detection
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 12f0d60f-a0a8-4de2-8900-5d89404a427b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Less is more: Understanding word-level textual adversarial attack via n-gram frequency descend
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 086aa09e-52d0-428c-9a82-d9f5c532ad10 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Training overhead ratio: A practical reliability metric for large language model training systems
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c5d159b-7570-46a2-934a-4c80b0ceaf2b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation baa0e1a6-2158-4c23-be02-da1a495353e6 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Training language models to follow instructions with human feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f38a6a-2757-427a-ac4d-a5f79d4df568 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Fine-tuning now available for gpt-4o, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c5cc1688-fbb6-4376-9516-d02af1e946e6 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Instruction tuning with GPT-4
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3409dd57-7632-4bc2-a924-08eab1087789 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1f14dfad-4f74-48a4-85cf-db3325442932 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets D., Ermon, S., and Finn, C
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3eff96cc-2f75-460d-a807-f4aff180614e · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Modelgrow: Continual text-to-video pre-training with model expansion and language understanding enhancement
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 50cf077a-867a-4fe9-96a3-029a9f7074d3 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Representation noising: A defence mechanism against harmful finetuning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation febc9c75-3f5d-47de-803c-9ce89aac5cfb · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9c59bfa1-d2d5-4804-8972-aa168d140938 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets do anything now
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8406427-e273-49eb-b614-0c6f8059cd1e · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Axiomatic attribution for deep networks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 91529f78-8d93-485c-9738-0e1c54cd2af4 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Intriguing properties of neural networks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a816b57d-fada-4e37-aa60-da5529643cde · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cfe80fa-a50d-485d-af6d-e2df5fa984af · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Llama 2: Open foundation and fine-tuned chat models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5469cc26-d445-4d2f-afe1-f8c12451fa4d · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4bbcd51a-bee2-4bd9-86d6-e7f46c1013ad · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets W., Lester, B., Du, N., Dai, A
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a1d094-073c-4a9d-947a-7c08c09d0c42 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ca1bf601-9710-4d28-a6e5-d615025475f9 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets T., and Zhang, Y
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e111a9b0-e93e-441d-9842-e85a989af69c · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Open the eyes of mpnn: Vision enhances mpnn in link prediction, 2025
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4352e134-1c60-455a-b689-5b751316a200 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Backdoor graph condensation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f0d71694-5392-408c-9660-75b5f1c6470b · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Tf-dcon: Leveraging large language models (llms) to empower training-free dataset condensation for content-based recommendation, 2025
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 00bc88b6-1acf-4493-aa6b-e2118a33cd38 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets RLCD: reinforcement learning from contrastive distillation for LM alignment
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1bc3c7b1-5acf-445b-969f-027fc208a9e9 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Y., Zhao, X., and Lin, D
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e73d58be-004b-41e2-8c42-1806d6eac4fb · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets NLSR: neuron-level safety realignment of large language models against harmful fine-tuning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7e91603c-a9ac-456e-83f9-cd977e6019ac · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c731b779-da08-4469-8551-bb95f636ea28 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets GPT-4 is too smart to be safe: Stealthy chat with llms via cipher
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 86e8a6a8-443d-4d40-beb8-24f16c2d4e51 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Removing RLHF protections in GPT-4 via fine-tuning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 648c4e5a-e75e-4bf7-82f3-634c0cf6e84d · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets P., Zhang, H., Gonzalez, J
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 70f4184b-93f8-4cff-b3e4-157072b49626 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Model tailor: Mitigating catastrophic forgetting in multi-modal large language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e0e6b709-8b36-48fe-a6fc-0777aea92182 · outbound
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets Z., and Fredrikson, M
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9df8334b-e98c-4196-b3b4-0d0ce55f9688 · inbound
Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link Prediction Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa548b6b-122a-4a6b-80d2-7e615c384983 · inbound
Continual Safety Alignment via Gradient-Based Sample Selection Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ad72020-6449-47aa-a97c-2bb7e0eae2e0 · inbound
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.