Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T05:21:30.132993Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2607.11475.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T05:21:30.132993Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c07110e2-2969-4442-868e-98d365be6268 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Fine-tuning aligned language models compromises safety, even when users do not intend to
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c928cdbd-4db0-4015-be5d-edff62e95a99 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Safety alignment should be made more than just a few tokens deep
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d563caf-1460-42b2-b935-43fafe2c5554 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7d8cdc-268f-44f9-a607-326b04d3700b · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Representation noising: A defence mechanism against harmful finetuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b47ad0e-5e76-48fb-be7c-dff62e15b6fb · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Safe LoRA: The silver lining of reducing safety risks when fine-tuning large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18c6944-241d-4d35-99be-70adbf265e91 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Safety at one shot: Patching fine-tuned LLMs with a single instance
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4516ea-b4d3-4376-9a15-5b469ba0de62 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Editing models with task arithmetic
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84fa6f04-1b9b-4a18-842e-e72fda1c11ba · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e0a8fc-8db3-4ed7-b23d-4e9e6490c449 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b783caa-c649-4cf4-8e97-3c668c3d182a · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Granite Guardian
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05191924-c973-4ad3-b526-e757b5634d00 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da82233-f7db-4e93-bec7-fcb4ca380c9b · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4db2e4-3a9c-49f8-801a-51f8daf003f0 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7eb379-fdd1-4629-9a64-a3ac2a4612ce · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Training language models to follow instructions with human feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9eca68e-e963-412d-b998-38e064ed5790 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03bbafa9-ee6d-405c-80b8-0280ee34ee0f · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Learning to summarize with human feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40583a79-7b9e-459d-aab6-16c3c6e5505b · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Deep reinforcement learning from human preferences
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877242e2-0df6-4f71-87b4-848580998d3e · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71524d68-05c1-4572-9b33-dd09b19e6a93 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ea002f3-ab18-4c3a-a54e-ad1cc87b4f0f · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f180efd-54ba-4654-9549-1e8b2accf6aa · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Qwen2 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d07ebb2-b144-496f-96f6-23ce61d27c18 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Assessing the brittleness of safety alignment via pruning and low-rank modifications
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74866dd-c3bb-439e-bd18-986f2c40858e · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Fraser, Hillary Dawkins, Isar Nejadgholi, and Svetlana Kiritchenko
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09230b7e-e5b8-4880-bf7e-c8d2c97b186c · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models LoRA fine-tuning efficiently undoes safety training in Llama 2-Chat 70B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4974b58c-f410-4f5c-a1b1-98ab6c6dd27d · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Removing RLHF protections in GPT-4 via fine-tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bee95b9-158f-44c2-a685-f6661945d8cc · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f714f8b9-5742-44f6-8617-3a6858393337 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models TIES-merging: Resolving interference when merging models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804b7d9e-0507-4857-bd08-4dfc01df6563 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Language models are homer simpson! safety re-alignment of fine-tuned language models through task arithmetic
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e69fe6d-d765-4a68-9824-9f27f0eb1975 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337424df-419c-42e3-bd92-b6f4bd8f8c1a · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d0fa2d-a920-49ff-a1d6-3683f8d3d691 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6efa4a77-78b3-4ed0-ab0b-afc0ca2985ff · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models A holistic approach to undesired content detection in the real world
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8e818a-43d9-485f-920f-fb654092d434 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models A new generation of Perspective API: Efficient multilingual character-level transformers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32e5c7e-39a4-46be-9a94-05b06f5f98d2 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models HiddenDetect: Detecting jailbreak attacks against multimodal large language models via monitoring hidden states
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8100717-1401-4962-a528-44796c5d841d · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models The internal state of an LLM knows when it’s lying
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcfc93d-377a-4c07-aca2-2927dae5ec29 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Refusal in language models is mediated by a single direction
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71825118-52a0-4184-af1f-37694b1f1fa4 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Representation Engineering: A Top-Down Approach to AI Transparency
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f54519-1a35-44b6-8409-86e37f567edb · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5beb94f-d4e3-44c4-8708-8589444dbd08 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models A brief review of hypernetworks in deep learning.Artificial Intelligence Review, 57(9):250, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ba8e0b-a414-4a1f-b949-a8e8741ab3a6 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Continual learning with hypernetworks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c67a91-971d-46d9-aa7e-108ee18e7cb9 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Equivariant architectures for learning in deep weight spaces
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4e1342-1992-408f-b54c-94da3765a9ed · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models HyperTuning: Toward adapting large language models without back-propagation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff055df4-24a2-4dfa-8cfd-dfbe632c75b2 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f994dd60-0eb9-4e85-a08c-24a4251896c6 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Parameter prediction for unseen deep architectures
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd87627-6cc1-4354-bbe1-e53f944fb518 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457e8930-1a63-4aa1-afd7-d96876c8e687 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models LST: Ladder side-tuning for parameter and memory efficient transfer learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adbd0c7-9aae-4642-965a-d307e059fb56 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models LoRA: Low-rank adaptation of large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50c7fb4-3f5f-4e92-86c8-fb336ea3f131 · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models BeaverTails: Towards improved safety alignment of LLM via a human-preference dataset
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43f8798-f7ac-4c67-909a-2569ba31694b · outbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.