Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2310.20624.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:23:02.390992Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:59:53.050459Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9316efe9-cb13-4923-b761-37516f01734b · inbound
Refusal in Language Models Is Mediated by a Single Direction LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6e7dbe39-3217-4a9a-8f83-2d7baaabdc48 · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 512ac141-c77c-4a46-bde9-1738150efb2a · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c79286d2-e8bf-4dca-bc5e-22518a51632a · inbound
Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001670e7-7c84-4315-bf67-bd245bd8a832 · inbound
Jailbreaking to Jailbreak LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad933c1b-7a01-4ba1-8ed3-cb314ddb6a47 · inbound
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · inbound
Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111021dd-05f6-465a-b48e-2376af7f3472 · inbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f57e83-21f8-4faa-998b-68177e063140 · inbound
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d9129f-7bb4-4215-ae7b-53ab8a841f8e · inbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f2add8-82a1-40b4-ae54-e40684a089a7 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · inbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0230405-d481-4232-a71e-a6bd91d0f21a · inbound
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3407591e-bc9f-4f12-bf67-eb608e0b06f3 · inbound
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39bbd99-254f-4599-8cfa-917f2b9d2688 · inbound
Understanding the Effects of Safety Unalignment on Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51b9dee9-9552-4825-aa11-e8f2e9bbb3a6 · inbound
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 422f1b45-3cb3-4390-9e7c-d6df5bc3ef48 · inbound
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da4a9885-859d-4dbf-8363-6f186afdc9e5 · inbound
Low-Rank Adaptation Redux for Large Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a03f853-c1c6-4675-89be-89d7d32986a4 · inbound
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0fc8591f-5942-4348-ae96-eb953144eb4b · inbound
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d233b8b-046f-482e-86fd-c6c11a5a36ed · inbound
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d80882a7-057f-4b7e-8b5e-9fb5cb69ae57 · inbound
Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 006257a1-00c8-49c0-bb49-ba8b76145c5f · inbound
Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f51e1717-a090-434d-a1fa-d80497739ad6 · inbound
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b52221c-a0a7-4743-9217-5e493c2088d3 · inbound
Adversarial Reframing: A Framework for Targeted Generation in Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a77825fa-713f-43ea-b688-d21710ce826e · inbound
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb7efac4-ad6d-4700-806d-b596f98a71b9 · inbound
Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 76c0ffc5-9fb8-420a-bd4e-450a48bf5004 · inbound
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e8bf7610-d2e3-4ecc-a0ea-df10c8c82bca · inbound
Curriculum Learning for Safety Alignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 671d674a-6cd5-42cc-b508-bc3daa08761e · inbound
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4fe3cc67-1891-4f66-ace7-8f8b60f7b0ee · inbound
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29586d0c-7565-48ab-9965-d676f03aa345 · inbound
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93a5ac7b-9a1f-40a6-af55-c8c43f8b5b4a · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1f12f49-34ff-415b-8778-47ce65e5716a · inbound
RepSelect: Robust LLM Unlearning via Representation Selectivity LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d16d39a-46d0-490c-bbfe-bfd4c1cd6bda · inbound
Discovering Millions of Interpretable Features with Sparse Autoencoders LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7763be8f-401c-46ff-834f-475b65181e5f · inbound
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 150f3f6e-8b64-4988-9ed7-8e975ea53f56 · inbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a946a932-8846-47e0-af95-2f39d7873f13 · inbound
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4e70e1-67fb-424a-937f-79b556eadb57 · inbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.