Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2311.09433.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:51.896624Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:27:29.192379Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3a09ebe2-8c9d-406f-af79-8900035ef1a4 · inbound
Refusal in Language Models Is Mediated by a Single Direction Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 42c003f5-cb69-47c0-9099-d328c904716e · inbound
AgentReview: Exploring Peer Review Dynamics with LLM Agents Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ebb851fd-c435-49bb-a464-ee806890944b · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63f11b28-45ac-4144-844b-1c0a5b001817 · inbound
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d93e93e-2304-431b-867b-c035a315f234 · inbound
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · inbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd41a068-d301-48c4-a177-4c4270a48e03 · inbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e77078-dd1c-41f3-9ab2-6c2cbbe4139d · inbound
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d030c5a8-2651-4432-b2c9-9454504a4bba · inbound
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbefdb4a-f310-4c2a-98cd-b930afcfbb91 · inbound
A Survey of Attacks on Large Language Models Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0edccbc-9058-4a21-ab17-e03d4e0de2a8 · inbound
Improving Multilingual Language Models by Aligning Representations through Steering Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15125eaa-6240-4c15-bf31-a7767070f8d6 · inbound
Security Concerns for Large Language Models: A Survey Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · inbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180e7a9a-bfa9-40af-a697-a47145e9befa · inbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · inbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4166b559-ecfe-405c-863f-59b3cf19d032 · inbound
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70146f0-2fba-4554-9a18-c148d61dcb7e · inbound
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 392ef71b-f54a-460c-9979-0a6a0328f509 · inbound
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669b354d-6352-46b5-ae79-5dd23e4f80a3 · inbound
Steering Protein Language Models Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df99dcde-4a01-4892-a680-119eedabcf21 · inbound
On the Privacy of LLMs: An Ablation Study Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78740697-7847-4edf-8364-1f2f8a477cf4 · inbound
Distilling Safe LLM Systems via Soft Prompts for On Device Settings Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ad231733-d94c-49a6-a5d4-45e21278f74c · inbound
Evading Chain-of-Thought Monitoring Through Model Poisoning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.