Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:34.839326Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.00676.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:34.839326Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de9f9c04-8478-4a35-b38b-bd27df643ce5 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427d424f-591a-4f31-9921-131da82106cf · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Zico Kolter, and Matt Fredrikson
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f90175f8-b759-464a-b357-e13422769c53 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Dai, and Quoc V Le
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df60811-3b14-484c-b131-d331a39ac21d · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training language models to follow instructions with human feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef84feb-83b5-440b-b385-26ebd177d1e8 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Direct preference optimization: Your language model is secretly a reward model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b992868-0ff3-44c0-b7b2-7794e294ae1a · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6e2760-ce4e-47db-9af0-a46a4b894e1a · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4aa1255-316b-49bd-9c0f-f8014a9a093d · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Gemini: A Family of Highly Capable Multimodal Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8713b269-1393-43b7-a7ca-55333114f870 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lora: Low-rank adaptation of large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f634a9af-af2a-4454-bc05-8e043d5ce904 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960ad58d-a2de-4693-9111-0cb30b55e1e7 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Pissa: Principal singular values and singular vectors adaptation of large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be6ac48-e3bd-41a8-bc43-8a79b5adbce5 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Badam: A memory efficient full parameter optimization method for large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f30c869-55e7-4b8e-95a4-2b6cf3286f8c · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Visual adversarial examples jailbreak aligned large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19e501af-73cf-4e17-9410-717948bfc92c · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89dc442d-4a46-413a-bad7-eec925eab060 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ff44d53-0ff6-4a37-a77c-cb02b38060c7 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 070c56b1-b024-4089-b0af-2b10a0bbeb88 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15a4efd1-f46c-4d61-8bf4-36c3999ba478 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4950454c-2309-4f59-b35d-9bdb1fb441c3 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Keeping LLMs aligned after fine-tuning: The crucial role of prompt templates
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05d67407-e3c4-4ece-a8dd-92fa0b9dec88 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safe loRA: The silver lining of reducing safety risks when finetuning large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6f6a29c-114a-4b38-8b5c-51670c8fb2c7 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Tamper-resistant safeguards for open-weight LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66edb223-d40c-4319-8cb9-b637a74fb4f7 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 532f6c18-5cea-401a-85e0-cf29774fab08 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety layers in aligned large language models: The key to LLM security
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33abe46c-c251-4133-85d3-d9ab564bb628 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SaloRA: Safety- alignment preserved low-rank adaptation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be3634a1-8495-4f10-b0bd-bc639ee0f10d · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32b908e5-7d3c-4586-a750-fd4df42093f9 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd8fa39e-d18d-4541-b2f9-f35aeb625878 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5920d55a-850f-44df-8058-206b8b396850 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety alignment should be made more than just a few tokens deep
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063617ce-60fb-4340-ac7a-71aaf5b3622b · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Defending against unforeseen failure modes with latent adversarial training, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 479a383f-5ceb-4520-ae5f-c557d2eecf89 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38cbbcc2-9445-4174-925c-36f6299ec164 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36b4a338-cdeb-46ba-8666-ca241a3f0cd2 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0220af5e-d502-47e3-8036-525afa5ab5d0 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98dead1d-25fa-4373-9a91-24834ff66963 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc923ced-3d93-49ee-b673-1ffdbd439a66 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Assessing the brittleness of safety alignment via pruning and low-rank modifications
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f151a94-4360-40f2-b2d0-2976dc036c4b · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Toward secure tuning: Mitigating security risks from instruction fine-tuning, 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f17451b-dd86-4a84-830c-b9a2252ea844 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Locking down the finetuned llms safety, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6d5ad28-eef6-4875-94b5-caa73e79424c · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training Verifiers to Solve Math Word Problems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2e8df3-365b-4eaa-b24b-c22f448e7c63 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Measuring massive multitask language understanding, 2021
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edab1b0-6c34-4396-9b13-872213a50348 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Xing, Hao Zhang, Joseph E
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbfd0b5-32a2-4dff-95fd-f293d9622854 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Certified defenses for data poisoning attacks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61c38025-b471-486b-b429-ff71b0240bfa · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Data poisoning attacks against federated learning systems
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c26f605a-8aa8-4e12-ad5c-382f47996d28 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Immunization against harmful fine-tuning attacks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a852dd-4a36-4ce8-a082-12588c67ed17 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ec8314-dc2d-4267-8e4f-5a0ad634537f · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning The case for an opinionated, theory-oriented real-time operating system
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a34b9a9-a42a-4873-81c9-41aee11e8d93 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c324371c-2ac5-4ddb-b0ec-8e653cc155b7 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Recursive deep models for semantic compositionality over a sentiment treebank
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca2bf94-8217-4af6-8aa8-00a1a4362ebb · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Character-level convolutional networks for text classification
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4746240-e2a8-4e3f-947b-886a93bb3f2f · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f838a6b-4c60-4c9b-90b8-d95bb9dde953 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4890fa07-bb75-4ae3-8670-a364756f90c1 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2893360-00ca-464f-b0f3-8d2aafcc35fe · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Alpaca: A strong, replicable instruction- following model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 463654fc-4125-455b-a554-7367748f69e6 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Introducing the world’s first truly open instruction-tuned llm
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e5921d9-feef-43d2-bd67-43c267879912 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning q_proj",
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 498903de-128a-47f0-8494-6b248ff36c3f · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning This is the most preferred response when confronted with policy-violating instructions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb12fda-29ea-44dc-a028-3c21b644ba91 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23d4b7a-4399-4401-96c0-e9c832dcc3c1 · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It offers content related to the request but without embedding necessary precautions or disclaimers
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b6152a-c0c0-4c9c-b8c3-eff51f401bda · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f610a2d-f3a6-44ba-99ce-8512d95179bc · outbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Candidate X Under Investigation for Embezzlement
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.