Pith. sign in

Paper Citation Record · LEDGER

Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2311.09433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09433 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:51.896624Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:29.192379Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a09ebe2-8c9d-406f-af79-8900035ef1a4 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.141141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:47b30965473b18a862c0857227330bf149c48d6d8e952dd04c5440049c52220a

Observation 42c003f5-cb69-47c0-9099-d328c904716e · inbound

AgentReview: Exploring Peer Review Dynamics with LLM Agents cites this paper.

AgentReview: Exploring Peer Review Dynamics with LLM Agents Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:38:37.121335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T23:38:28.005028Z digest=sha256:4280b473832dd02011ba1a009818b93691351584387be2251167193e61b50d43

Observation ebb851fd-c435-49bb-a464-ee806890944b · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.320120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:5944d75408cbeb389445c7aa669cd87ff892be1ce2f7fb2bbfd041fe88188528

Observation 63f11b28-45ac-4144-844b-1c0a5b001817 · inbound

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations cites this paper.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.932852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.932852Z digest=sha256:5eb493a8ce5ec23f19af9bf7b70ab4cf436b798b8fd59380aeb543360881767b

Observation 2d93e93e-2304-431b-867b-c035a315f234 · inbound

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs cites this paper.

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:19.111201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:19.111201Z digest=sha256:b0875e4529f26e746f0e0fda380c129643e80acaffa3c40f0d021632f502fdd1

Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.848329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.848329Z digest=sha256:a639263a8b80ade7da768a253ba82e38c1fec793e8952aae0d9124a24d2315e9

Observation dd41a068-d301-48c4-a177-4c4270a48e03 · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.367750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.367750Z digest=sha256:342701fb7ae7b33b167f992fa98d229c88ba0e12e409843e361d089bc487cecb

Observation 91e77078-dd1c-41f3-9ab2-6c2cbbe4139d · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.594555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.594555Z digest=sha256:36d0e6163fdbe33f567f3e67bef6603d6392ef0fbe95925da3d135e39cb5729e

Observation d030c5a8-2651-4432-b2c9-9454504a4bba · inbound

BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts cites this paper.

BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:51.896624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:38:51.896624Z digest=sha256:861de4d6e55d4f74ad1fefeecf27324b8849caa1d4747e3fb5e00566e1597196

Observation fbefdb4a-f310-4c2a-98cd-b930afcfbb91 · inbound

A Survey of Attacks on Large Language Models cites this paper.

A Survey of Attacks on Large Language Models Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:34.367090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:34.367090Z digest=sha256:3a3fd854c33f2b308da5e92fc146af4efa00977072de094dbd57312cdf84f9ae

Observation e0edccbc-9058-4a21-ab17-e03d4e0de2a8 · inbound

Improving Multilingual Language Models by Aligning Representations through Steering cites this paper.

Improving Multilingual Language Models by Aligning Representations through Steering Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:27.342800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:42:27.342800Z digest=sha256:3ea977ed11562c25e4d3106dc2645aa1dc117852bf92f765da8d73db599cf9a3

Observation 15125eaa-6240-4c15-bf31-a7767070f8d6 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.694102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.694102Z digest=sha256:1ec9bacfbabacad37667e5eac21e79441ba0e6eee3bf31e4b61c521322438a8b

Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · inbound

Probing the Robustness of Large Language Models Safety to Latent Perturbations cites this paper.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.484246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.484246Z digest=sha256:1579c3bf8eb5752956b967bd9d581070e4ceaaca476b15a7b1b10af33ad9ec08

Observation 180e7a9a-bfa9-40af-a697-a47145e9befa · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.044270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.044270Z digest=sha256:ddf4b46314912ec76a63ba3f9f9f7726678818d36d110fd3546c2df07a5b344e

Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.929314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.929314Z digest=sha256:e3d5e10aaf6a30dc48a94ed19fe497dc182d08ed1bf54ce6d3763c2ad610c539

Observation 4166b559-ecfe-405c-863f-59b3cf19d032 · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:28.669109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:28.669109Z digest=sha256:f33cf444c8361decf85a84715bd88b8fbd70ced325c674c6c5861888e5faa4e0

Observation c70146f0-2fba-4554-9a18-c148d61dcb7e · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:28.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:28.741450Z digest=sha256:c712c84f4924f6ef0231a23c080d6cd1a4afbe391a747ddb9e1c92f72f7ab812

Observation 392ef71b-f54a-460c-9979-0a6a0328f509 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:23.762799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:23.762799Z digest=sha256:734bbc34e0367b46fec0a18708f26353180e7ded83406aa9b4157ddbb5d05635

Observation 669b354d-6352-46b5-ae79-5dd23e4f80a3 · inbound

Steering Protein Language Models cites this paper.

Steering Protein Language Models Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:12.120075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:12.120075Z digest=sha256:58f85911473280182da42fda5d422eacdb139e689eb70532b874991ffca54014

Observation df99dcde-4a01-4892-a680-119eedabcf21 · inbound

On the Privacy of LLMs: An Ablation Study cites this paper.

On the Privacy of LLMs: An Ablation Study Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:48.822645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:25:05.586464Z digest=sha256:bd81f3c3eececcb910bae740b0e685e47a519e156fbf7b888e863697e22973c2

Observation 78740697-7847-4edf-8364-1f2f8a477cf4 · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.194109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:5d622c6185984e81c1ead38738bfdd0a09dda23949d6f26e4aee93581a7f214e

Observation ad231733-d94c-49a6-a5d4-45e21278f74c · inbound

Evading Chain-of-Thought Monitoring Through Model Poisoning cites this paper.

Evading Chain-of-Thought Monitoring Through Model Poisoning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:04:48.284542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:04:48.284542Z digest=sha256:f18f04e7501946ee2786195522a5654c631386bf8c2d28ba91b5e534145ce31e