Pith. sign in

Paper Citation Record · LEDGER

Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.09433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09433 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:03.694102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:29.192379Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a09ebe2-8c9d-406f-af79-8900035ef1a4 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.141141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:87645c8c992ba75635e4c3463cb4cac017f9bc85a4a0633fe1473b7fa07c88c0

Observation 42c003f5-cb69-47c0-9099-d328c904716e · inbound

AgentReview: Exploring Peer Review Dynamics with LLM Agents cites this paper.

AgentReview: Exploring Peer Review Dynamics with LLM Agents Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:38:37.121335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T23:38:28.005028Z digest=sha256:3ce35174aa928be08f8f0611c062a38a8e4839378c9bc3ed2f800b2802f98b4f

Observation ebb851fd-c435-49bb-a464-ee806890944b · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.320120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:2f5ddb8b750c66ef48c809c8e0fd0e7aa7a1a371b4d1233fbf5f38107c88a3e5

Observation 15125eaa-6240-4c15-bf31-a7767070f8d6 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.694102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.694102Z digest=sha256:8d7db825876518a1d557b51aba3316fb755fb812a21778ab05cb30b26c32a739

Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · inbound

Probing the Robustness of Large Language Models Safety to Latent Perturbations cites this paper.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.484246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.484246Z digest=sha256:e92138b5669e0ce00ca4313fed74e32dea93190231bb32bf2e76518b6e8b7616

Observation 180e7a9a-bfa9-40af-a697-a47145e9befa · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.044270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.044270Z digest=sha256:603091826e831c74be5b2c337c69e266cfbd077b825f2db29da5761efd4aa79f

Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.929314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.929314Z digest=sha256:0093ff15901cf3ba906dbcbc119ea73ccff75e14c09d4d7acb527eeba54c7b99

Observation 4166b559-ecfe-405c-863f-59b3cf19d032 · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:28.669109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:28.669109Z digest=sha256:e15225f0651baf57d2ea6db21722b56a0d28aa18ae6a3b1a4b4d8c0a6a945069

Observation c70146f0-2fba-4554-9a18-c148d61dcb7e · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:28.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:28.741450Z digest=sha256:a53d84805b0d85ffe94818db4d9b4d9dbccf89279a269953f1e5f5f38b08d083

Observation 392ef71b-f54a-460c-9979-0a6a0328f509 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:23.762799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:23.762799Z digest=sha256:73adcdc959040013114f77f56733dfd688bcdcce4c74e3508b278ce1c86d987d

Observation 669b354d-6352-46b5-ae79-5dd23e4f80a3 · inbound

Steering Protein Language Models cites this paper.

Steering Protein Language Models Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:12.120075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:12.120075Z digest=sha256:29411db95826f93a4310d68a878d1ec2486d29e3999e5725ab13b61b96c3bceb

Observation df99dcde-4a01-4892-a680-119eedabcf21 · inbound

On the Privacy of LLMs: An Ablation Study cites this paper.

On the Privacy of LLMs: An Ablation Study Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:48.822645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:25:05.586464Z digest=sha256:ca90560c05ee20bd95a445737cc89c997dd61764931fbc5784ea13f9af840602

Observation 78740697-7847-4edf-8364-1f2f8a477cf4 · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.194109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:c5be9db8a737060a4c67cf47a47f95dfb39cb8294584d28cfee963ab682ad812