Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:12.986775Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.20766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:12.986775Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2aa4da68-6b4f-4008-b654-dcbe48a3917c · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Yi: Open Foundation Models by 01.AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092afef3-3837-4b5e-8f95-9b1ae6b685bb · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal in Language Models Is Mediated by a Single Direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff2285ec-1da2-4be1-9a5b-cdc9c5650f98 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468cad06-0e4a-414a-8c22-166bb21232c4 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards inference-time category-wise safety steering for large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8333e408-0fca-4c09-8e88-c446a4ecf603 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e117786-9786-498d-b6c0-fdacfa42fa86 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards monosemanticity: Decomposing language models with dictionary learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e38398-c650-48ed-bb6e-9c72e74f13cd · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17474615-b21a-4117-bbc4-ac63ae9ab4fa · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On the Measure of Intelligence
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd21b06-c7ea-41e7-b016-ad9b2b1c5695 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b00bbb-6d0c-4ff4-b25f-da59a5d282f7 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Boolq: Exploring the surprising difficulty of natural yes/no questions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f2370b-d7be-4c1f-aeff-02c1b5730df9 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection https://dphn.ai, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c782017-4040-435c-952c-0013deccf568 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Toy models of superposition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12d6ed4-a3ec-41f5-b817-e37bfd8aa8aa · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Finding alignments between interpretable causal variables and distributed neural representations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae0c54c9-a499-4401-81f9-d061c0c906b9 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2f0ac79-d72b-4a3a-b0fe-910bf950ab7e · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4271ee28-0264-4c95-8ded-ef9226004dbb · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Model merging and safety alignment: One bad model spoils the bunch
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b3b0e0-4190-4f7b-812b-b3ae8e6a6549 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad500bd-6164-4250-bb38-159e43dab79b · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e8ee80-4730-41b1-b93b-9511a0285653 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca596ab6-2639-4d8b-8f7c-4b3ed707e016 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Measuring Massive Multitask Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d963b07-a33a-4bee-b4dd-929f49162067 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a161f0dc-f6f2-4fcf-8c5e-44d44cfe8f82 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 302948ed-b1c0-45e3-b42f-698298bbc876 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection What makes and breaks safety fine-tuning? a mechanistic study
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99bad9e1-902a-4446-b2a1-a415beaa3217 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc4d333-17c4-4fbc-a229-329185933c1b · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Evaluating Open-Domain Question Answering in the Era of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9662f8-177f-4f6a-9e85-d35648a8fbd0 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Inference-time intervention: Eliciting truthful answers from a language model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb2e00b-e697-449a-a39c-6d85ae6cf5ae · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Rethinking jailbreaking through the lens of representation engineering, 2024 b
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3844fb-c2cb-4f88-91fb-61e375962b98 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6023b781-e395-45a4-8076-f3037ddade19 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards understanding jailbreak attacks in LLM s: A representation space analysis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28e1915-567e-4f34-87db-0b7f74f29fb1 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Llama 3 Herd of Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486f20dd-9218-40a3-abae-e940bb414b97 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f9c9c6-d705-4f4e-baa3-cb5f4dfe2c17 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb034246-6af4-4a59-b1e4-b960e64aae42 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2db5bc-92c8-4a3e-aba5-1c89c5bf2c6d · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Model Refusal with Sparse Autoencoders
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486399a8-beeb-467e-bcff-766c87154cd1 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Training language models to follow instructions with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd42cd7-bdda-4659-9abe-4c29ecee0864 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Llama 2 via Contrastive Activation Addition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57a135cf-df8c-4683-85cc-9bd623cf73d4 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0671c2ed-033c-458b-b1c9-eae1ae50a15c · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Qwen2.5 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea3a346-6727-4f0e-a0ac-fe9f4cf288e2 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec3c9fe-56bb-47cb-8083-252dbc476f9b · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection An embarrassingly simple defense against llm abliteration attacks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97459c1a-bf2f-4334-9e57-cff5bdd29dd2 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Hashimoto
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92c241b-2aba-4afe-b9b1-ae0890ac495e · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Gemma: Open Models Based on Gemini Research and Technology
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdda36b2-8d12-4090-9c2c-91c91c5d6d0d · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Daniel Freeman, Theodore R
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99dd246e-57ef-45d8-b199-ff87eda4e38a · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection CodeJudge: Evaluating Code Generation with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b876ede-e860-47df-a75a-7f728f26ebd3 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab913e3-618f-44bc-8a36-de6a9893441d · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Models With Activation Engineering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9949d7-4e83-4504-be8a-a821d8024bfc · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 853dddd6-f99c-4b48-a081-8870c95abc88 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Jailbroken: How Does LLM Safety Training Fail?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bd386f-8001-44d2-b4cf-dc27d23048cc · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ab188a-2c4f-45f4-b452-5e3f2a36c52b · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d8626b-c077-46ae-b78c-1ecdbeb34004 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e8c9d3c-aa7e-4eb7-9ed2-fbbabcde399c · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation bending for large language model safety
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50435e64-1521-4101-838f-da39347e0c24 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Robust LLM safeguarding via refusal feature adversarial training
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d65efbe-17ed-44bd-b2c3-53e3c656be66 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d37d7c29-99b8-41d1-a143-1a3e66ed316f · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Adasteer: Your aligned llm is inherently an adaptive jailbreak defender
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96e3b12-54fb-4287-967a-e210244f832d · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On Prompt-Driven Safeguarding for Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf27ccbb-f785-4237-bce1-f84d60fdb599 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation Engineering: A Top-Down Approach to AI Transparency
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2082d8e4-ab00-4930-86be-38c92635280c · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection write newline
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9341250d-b03c-4094-a3f2-8a4ebbbe4621 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection @esa (Ref
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47659d76-c0b0-4841-aa01-f10e23eebb01 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1294e0a-e785-4981-bb46-198da192b190 · outbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.