Pith. sign in

Paper Citation Record · LEDGER

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.20766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20766 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:12.986775Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2aa4da68-6b4f-4008-b654-dcbe48a3917c · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Yi: Open Foundation Models by 01.AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.761630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.761630Z digest=sha256:5f104ee4aea9174b334eefd5aa9149bcf00c6ac1d4b13ebcb501a7bdaac48a2a

Observation 092afef3-3837-4b5e-8f95-9b1ae6b685bb · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.766265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.766265Z digest=sha256:88cafce16f5392182a4248653ce37a13f73feeeab8f10f947eb9af63aa0d6237

Observation ff2285ec-1da2-4be1-9a5b-cdc9c5650f98 · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.770432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.770432Z digest=sha256:be95d8713c259f28a9b33d81ffce51ac554576d210ba9ed845f79cfdb3ccfeba

Observation 468cad06-0e4a-414a-8c22-166bb21232c4 · outbound

This paper cites Towards inference-time category-wise safety steering for large language models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards inference-time category-wise safety steering for large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.899160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.774158Z digest=sha256:bf61229b22e5e64f083fbc5c2d977baef15cecbd66fceab796aa4d2996ad02a4

Observation 8333e408-0fca-4c09-8e88-c446a4ecf603 · outbound

This paper cites Man is to computer programmer as woman is to homemaker? Debiasing word embeddings.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Man is to computer programmer as woman is to homemaker? Debiasing word embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.777705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.777705Z digest=sha256:b527ec774119dd595b87731a351b37746e3bf54d11d86fd7112dad029f99c644

Observation 1e117786-9786-498d-b6c0-fdacfa42fa86 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards monosemanticity: Decomposing language models with dictionary learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.781141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.781141Z digest=sha256:7c5354b39ac122ba380509e55197903d15773118d192ba542acef49d999e3d2a

Observation 35e38398-c650-48ed-bb6e-9c72e74f13cd · outbound

This paper cites Language Models are Few-Shot Learners.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.784667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.784667Z digest=sha256:485bc1c9416085380148a96ac63b1773a024d60050e26dafbdc9fece7d1c4c5d

Observation 17474615-b21a-4117-bbc4-ac63ae9ab4fa · outbound

This paper cites On the Measure of Intelligence.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On the Measure of Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.788366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.788366Z digest=sha256:db981c5bc2f98b6c1baa9a016347ac4f5d54bb1a5cdf74701faf2e27d14e7d08

Observation 4cd21b06-c7ea-41e7-b016-ad9b2b1c5695 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.791829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.791829Z digest=sha256:c69343bafe20f222c2a9f00e9da8bd1144e018564b8ff8b79e05d59570d2c3ca

Observation 66b00bbb-6d0c-4ff4-b25f-da59a5d282f7 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.795167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.795167Z digest=sha256:1de59814c4589da3193585efba646bcc4dd08b0f1bbee8536eb68d00d8e5d5fa

Observation 51f2370b-d7be-4c1f-aeff-02c1b5730df9 · outbound

This paper cites https://dphn.ai, 2025.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection https://dphn.ai, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.872091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.798641Z digest=sha256:32c5bb34ec1832226ea43be501c78f9e99aac1ff7c074949e4ce2f291e489351

Observation 2c782017-4040-435c-952c-0013deccf568 · outbound

This paper cites Toy models of superposition.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Toy models of superposition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.801847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.801847Z digest=sha256:60fa8ffe58ee5ffde7f24b4252a9656b4e50d4b0d9c93fce157f42f98f0ef150

Observation a12d6ed4-a3ec-41f5-b817-e37bfd8aa8aa · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Finding alignments between interpretable causal variables and distributed neural representations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.856147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.804855Z digest=sha256:bd821d3821114283042789ff904f7c9b79e47c3a1562551cc745a59c0ce8a743

Observation ae0c54c9-a499-4401-81f9-d061c0c906b9 · outbound

This paper cites SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:57:13.715920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.808005Z digest=sha256:bb68802acb830ee6febdd34751e5885f2304bfd479a38d879d23db4d30a995a7

Observation a2f0ac79-d72b-4a3a-b0fe-910bf950ab7e · outbound

This paper cites A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.811315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.811315Z digest=sha256:c7bcb0b9897378d30f0429001bd5a26afa4a79af923fc89e013073fdde84b6e6

Observation 4271ee28-0264-4c95-8ded-ef9226004dbb · outbound

This paper cites Model merging and safety alignment: One bad model spoils the bunch.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Model merging and safety alignment: One bad model spoils the bunch

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.814839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.814839Z digest=sha256:eb79dfe6eac84a2d5ca26478ef05129d8b96bcd479e4929e4cc10307e3e023df

Observation 47b3b0e0-4190-4f7b-812b-b3ae8e6a6549 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.818090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.818090Z digest=sha256:0f5d99148a9e8daca2a7bae6edfa56754feb56911035833df7a550bb82c00e2e

Observation aad500bd-6164-4250-bb38-159e43dab79b · outbound

This paper cites Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.821622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.821622Z digest=sha256:53a2686d9d111dac1ec652dd302573e990b991b14af47efcfc510fc691e667e3

Observation 33e8ee80-4730-41b1-b93b-9511a0285653 · outbound

This paper cites SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.825477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.825477Z digest=sha256:11f812c0675bea67dac0bd36e78b871c9c6c67dd24724261eff4efd707618381

Observation ca596ab6-2639-4d8b-8f7c-4b3ed707e016 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.829123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.829123Z digest=sha256:c1272c3d89ec6acc45f8de23bce5c025c9b058e61fe6c3e879091a4bf6058d08

Observation 1d963b07-a33a-4bee-b4dd-929f49162067 · outbound

This paper cites The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.832694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.832694Z digest=sha256:41d66b874f3bc5f0971707aef5f6714ffd9e267dc044a7c67c62cb2dd2d21b62

Observation a161f0dc-f6f2-4fcf-8c5e-44d44cfe8f82 · outbound

This paper cites Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.836202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.836202Z digest=sha256:b170b75587a652b6c7eb45b3649519a7d5db2b5346d4be587dfd9424bf85a00d

Observation 302948ed-b1c0-45e3-b42f-698298bbc876 · outbound

This paper cites What makes and breaks safety fine-tuning? a mechanistic study.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection What makes and breaks safety fine-tuning? a mechanistic study

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.845319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.839957Z digest=sha256:1b54be281d6befb0c670b2c589f043fbfc7ffa8f0e59f004641f2a3358b62276

Observation 99bad9e1-902a-4446-b2a1-a415beaa3217 · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.843060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.843060Z digest=sha256:562a654240878bc5aef92a5518aa39eae0079eee6a92908078a6302ebbbe2042

Observation 0fc4d333-17c4-4fbc-a229-329185933c1b · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.846665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.846665Z digest=sha256:dfa3535c98ec1203b2bd2c54bf8e6add153d2f88432401c9ff41e741320ca1a8

Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.850194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.850194Z digest=sha256:64ca351d481b4459837cb4d05be9ab46c7fbfec90dc35177b24faba1699c12c4

Observation 6d9662f8-177f-4f6a-9e85-d35648a8fbd0 · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Inference-time intervention: Eliciting truthful answers from a language model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.854241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.854241Z digest=sha256:8fa56a8dfad8c554a44dcabd1ddef407930072c2d6c36aa6a7d700bae45687d6

Observation cfb2e00b-e697-449a-a39c-6d85ae6cf5ae · outbound

This paper cites Rethinking jailbreaking through the lens of representation engineering, 2024 b.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Rethinking jailbreaking through the lens of representation engineering, 2024 b

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.857443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.857443Z digest=sha256:b154a4c487b47ec18a53f1b9c113a82d13b3704212da9813bd4d750472219c9b

Observation 1d3844fb-c2cb-4f88-91fb-61e375962b98 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.860810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.860810Z digest=sha256:28dcee1b90ff7e9f2afa808e6aba9731639d74e0ed70321538830b2b074873b4

Observation 6023b781-e395-45a4-8076-f3037ddade19 · outbound

This paper cites Towards understanding jailbreak attacks in LLM s: A representation space analysis.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards understanding jailbreak attacks in LLM s: A representation space analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.864421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.864421Z digest=sha256:3b59391a09990ba1712d0c298fab6350ef419d1fb5b77056db0e9a4610c0f428

Observation d28e1915-567e-4f34-87db-0b7f74f29fb1 · outbound

This paper cites The Llama 3 Herd of Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.867614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.867614Z digest=sha256:37d43d2dd0ee647218e14fd4bc4b443db62e31e0cc8b30f2fa566737dcec73d2

Observation 486f20dd-9218-40a3-abae-e940bb414b97 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.871465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.871465Z digest=sha256:dff1ad4d2f0478f31333bcc4ee7bbb2191f52fc0464fa355876337b994b77865

Observation a6f9c9c6-d705-4f4e-baa3-cb5f4dfe2c17 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.875001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.875001Z digest=sha256:6f14380a633be01b387ee7dfb56d2072b7cb24e4fce6477557859075985a56cf

Observation bb034246-6af4-4a59-b1e4-b960e64aae42 · outbound

This paper cites Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.878381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.878381Z digest=sha256:b254877e037f3ffd2343c011e9d9b3f486d645daaa2449ecd013859ec8e90762

Observation cf2db5bc-92c8-4a3e-aba5-1c89c5bf2c6d · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Model Refusal with Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.882025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.882025Z digest=sha256:7c3f500eb1eed18faacae22216724d0eefd9a873cea9e36e9e2bcea73b76c439

Observation 486399a8-beeb-467e-bcff-766c87154cd1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.885402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.885402Z digest=sha256:6f47a936b5f3de48ea033c2401501a649d038c064a7ae0df266abbb518d96c39

Observation acd42cd7-bdda-4659-9abe-4c29ecee0864 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Llama 2 via Contrastive Activation Addition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.889009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.889009Z digest=sha256:f9cab47f83d87c7705869d99d00b2832891588d16042d67e465781ea3994cb89

Observation 57a135cf-df8c-4683-85cc-9bd623cf73d4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.892755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.892755Z digest=sha256:7bd64cee0d8c461fb9e7c5edd92449cef240ac2ca13d307ce0655fd2f1eb7684

Observation 0671c2ed-033c-458b-b1c9-eae1ae50a15c · outbound

This paper cites Qwen2.5 Technical Report.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.896258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.896258Z digest=sha256:cf16d0498ce76d896771a4729ecffcf4d0245bbd1b1c21dbb4bf6baf94b57e4f

Observation 2ea3a346-6727-4f0e-a0ac-fe9f4cf288e2 · outbound

This paper cites Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.900075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.900075Z digest=sha256:7f57e3ea0762d4fdcc0b1f7f2153011b7f1b310485cc653f2c46e70c2389f700

Observation fec3c9fe-56bb-47cb-8083-252dbc476f9b · outbound

This paper cites An embarrassingly simple defense against llm abliteration attacks.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection An embarrassingly simple defense against llm abliteration attacks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.903476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.903476Z digest=sha256:d5a6b11df5a2cf101a15aeab5fe94238db8ab9603637ae679d2011e9d9cd2c5f

Observation 97459c1a-bf2f-4334-9e57-cff5bdd29dd2 · outbound

This paper cites Hashimoto.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Hashimoto

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.906779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.906779Z digest=sha256:cf78f7714f94817add16b66fd5bc93922db2a2fbb5faac6215d9a5c923e593af

Observation d92c241b-2aba-4afe-b9b1-ae0890ac495e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Gemma: Open Models Based on Gemini Research and Technology

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.911134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.911134Z digest=sha256:eea645967f4385726941e4a21a48e5bd532e2f51961dafe8578aa2316c125bd7

Observation bdda36b2-8d12-4090-9c2c-91c91c5d6d0d · outbound

This paper cites Daniel Freeman, Theodore R.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Daniel Freeman, Theodore R

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.915277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.915277Z digest=sha256:b293b08903c9e8706e3f3e9124615d01b7664381a9f6afe676d2567801d72046

Observation 99dd246e-57ef-45d8-b199-ff87eda4e38a · outbound

This paper cites CodeJudge: Evaluating Code Generation with Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection CodeJudge: Evaluating Code Generation with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.918674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.918674Z digest=sha256:4b0c56323ae12c3c478debcfef618ccbef8ac501641a9f87455e0590a37a65e2

Observation 8b876ede-e860-47df-a75a-7f728f26ebd3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.922406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.922406Z digest=sha256:3ba6581b73fe2b8af96bb769921a070c0a4f2d66d2a26e2ec69f8e3adb65d61b

Observation bab913e3-618f-44bc-8a36-de6a9893441d · outbound

This paper cites Steering Language Models With Activation Engineering.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Models With Activation Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.926010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.926010Z digest=sha256:333995f558d7b04c5a516569b2fb9bb1b2821ce0caa25cbe015a0b946821c259

Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.929314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.929314Z digest=sha256:0093ff15901cf3ba906dbcbc119ea73ccff75e14c09d4d7acb527eeba54c7b99

Observation 6f9949d7-4e83-4504-be8a-a821d8024bfc · outbound

This paper cites Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.810919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.932659Z digest=sha256:9f388f3bdd69a3cd15ccca0860df0794eb318ab536ee9b75b7d97994979ee047

Observation 853dddd6-f99c-4b48-a081-8870c95abc88 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Jailbroken: How Does LLM Safety Training Fail?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.935639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.935639Z digest=sha256:11b6aea41e12a6306e2130c593173d111db27c8b337bdc3c6730ba924e978537

Observation 83bd386f-8001-44d2-b4cf-dc27d23048cc · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.939155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.939155Z digest=sha256:f17d88bc5542a18b2fc7e4f2a4d741adbd8ed9ee5acfe6f338c7b7449d5187ae

Observation b9ab188a-2c4f-45f4-b452-5e3f2a36c52b · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.942505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.942505Z digest=sha256:b8f8749767d134d42597689cfcc36de80b217a34b891fd5e0481283ce324c8a7

Observation 55d8626b-c077-46ae-b78c-1ecdbeb34004 · outbound

This paper cites A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.945857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.945857Z digest=sha256:612dba18750b1753f2a19165777e3758a1cda960cd551636213ab394bad5cca3

Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.949136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.949136Z digest=sha256:18a40fedd92819c576f5efdaf8150fc14493aafebec354301ed1ec3676a1a993

Observation 2e8c9d3c-aa7e-4eb7-9ed2-fbbabcde399c · outbound

This paper cites Representation bending for large language model safety.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation bending for large language model safety

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.952574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.952574Z digest=sha256:58749bed24d4985dd908d79f56abd39cd8062258228210a832d46bffb6fe3dc6

Observation 50435e64-1521-4101-838f-da39347e0c24 · outbound

This paper cites Robust LLM safeguarding via refusal feature adversarial training.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Robust LLM safeguarding via refusal feature adversarial training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.955790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.955790Z digest=sha256:b9ac999ae014d277426237f74dd36b461b482e2b41c9bddc8d113e42b769e155

Observation 6d65efbe-17ed-44bd-b2c3-53e3c656be66 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.959186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.959186Z digest=sha256:a5a5181867cadf36712cf3c422987ceee7e00f586aa8a90346598fb9d0e0da10

Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.962460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.962460Z digest=sha256:67698e0583f06eaec95b4769cf12c3d8cb5e4cd15156015448217005f0be2633

Observation d37d7c29-99b8-41d1-a143-1a3e66ed316f · outbound

This paper cites Adasteer: Your aligned llm is inherently an adaptive jailbreak defender.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Adasteer: Your aligned llm is inherently an adaptive jailbreak defender

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.965956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.965956Z digest=sha256:be59eb7b15c3aa2dba598f83fcdb98b17d0b13157eff18f276006aeb76b0ab14

Observation c96e3b12-54fb-4287-967a-e210244f832d · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On Prompt-Driven Safeguarding for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.969212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.969212Z digest=sha256:19b4e19ab66487e030409cded6064012f5cf1833bb57ac445c0c7c8ac94405be

Observation bf27ccbb-f785-4237-bce1-f84d60fdb599 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation Engineering: A Top-Down Approach to AI Transparency

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.972601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.972601Z digest=sha256:57170519de66b321d7c207a03679cef1cc7849ce8529d776cd309994eb476736

Observation 2082d8e4-ab00-4930-86be-38c92635280c · outbound

This paper cites write newline.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.976053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.976053Z digest=sha256:503b14c4bb6c7fef6a5788f273419c110a7e95fe12bf766826b3e554b9ebca0c

Observation 9341250d-b03c-4094-a3f2-8a4ebbbe4621 · outbound

This paper cites @esa (Ref.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.979904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.979904Z digest=sha256:8274a4b490c57b31ff364c0038a9232ff77ebc415395f3de3a8c3161a1fd15a5

Observation 47659d76-c0b0-4841-aa01-f10e23eebb01 · outbound

This paper cites an unresolved cited work.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.983359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.983359Z digest=sha256:7f974d3ea2a6e701fcbe5dcd6ce8f1b465112d8cf52da237b48b711f463a8bf0

Observation a1294e0a-e785-4981-bb46-198da192b190 · outbound

This paper cites an unresolved cited work.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.986775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.986775Z digest=sha256:73a011f9513753a2d5c6edcfbbdec7d0d5e1fdd1f80b51ec662f878d8acc9d10

Pith citing papers

No inbound Pith citation observations are available.