Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2310.02949.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:32.316238Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:09:43.659817Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 52fff216-cd20-4ab6-bdac-3a7b0be0fd59 · inbound
LLM Agents can Autonomously Exploit One-day Vulnerabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7a14b8b-4a05-474c-aec7-8cd21659b5b8 · inbound
Refusal in Language Models Is Mediated by a Single Direction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 202
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b20987fe-9db7-4022-a645-8d8002f9c10b · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1107c248-3eca-47d0-ada5-8b11aaa71381 · inbound
Learning to Ask: When LLM Agents Meet Unclear Instruction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e7154fa9-a96d-40ee-9d84-d5124b36211b · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · inbound
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574b4f9d-c622-4a9a-873a-34bff906c65b · inbound
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8423a4ca-4720-4e34-9c9a-ca8ba783decd · inbound
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916ad6f6-f012-41d2-8414-a9fbb17f21ee · inbound
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c2c932-e692-4384-ad68-e40230d919db · inbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c208afa-7b3a-4326-84bd-828646d8ed1c · inbound
S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · inbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1737b1d-a775-4801-93c3-1501881e726a · inbound
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4b329b83-cbf4-4dc4-b7f5-9c8b8ee0aaee · inbound
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9274f2a-95f3-4387-aa55-e4ea777a042f · inbound
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3597325-4084-4808-804f-767418eaaee7 · inbound
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a2df69-b7c7-4aa0-bc5f-0db29504e29f · inbound
Robust Policy Optimization to Prevent Catastrophic Forgetting Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 539c423a-598c-47bd-a544-9f2cdb64d3e4 · inbound
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a20819c-bd6d-4dac-97c7-d940b96d1f11 · inbound
Continual Safety Alignment via Gradient-Based Sample Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3982e68c-4587-4aa8-b7ef-d5b29f6f0bb2 · inbound
Representation-Guided Parameter-Efficient LLM Unlearning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac357e15-b452-4ffa-9014-d6ea0e5def20 · inbound
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c007f23-3612-4364-9134-03e5d47b5c66 · inbound
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 026f3de0-b29c-4319-8c5b-8b46131ae44a · inbound
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b05f4bd-9405-42b9-bc68-ec7f5a25c467 · inbound
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca6d5dbd-88cc-414b-a563-b525b536752d · inbound
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 363a98b2-587f-40db-bea0-4c105b6ccaa9 · inbound
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2e55edf-6079-4b2a-bbfc-0111bee5acf3 · inbound
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b84fdc6-85e8-4316-85e1-129d7639b7ab · inbound
Adversarial Reframing: A Framework for Targeted Generation in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation edb54687-7a19-4cea-a398-f22302efe7ad · inbound
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7cacd102-5732-4aed-8bc9-feaef39b7a1c · inbound
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42583a00-abd0-4214-a5eb-7b6dfc4c5c72 · inbound
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86a45327-cef7-4f8f-bcae-13b1650f33df · inbound
Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0ffbb66-00a2-4a1e-8735-f2177e4e3f6b · inbound
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bf2da3a5-6d73-4855-ad5b-43ac28e24aad · inbound
CSULoRA: Closest Safe Update Low-Rank Adaptation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 916ab45f-0e62-42bf-937a-b2dac39d344a · inbound
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c1a7d728-736d-4eb8-b6dc-091baa690fe3 · inbound
CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d5f89a3-ee3c-4a53-9222-020f6c0cd61a · inbound
Jailbreaking Multimodal Large Language Models using Multi-Clip Video Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7687dd7-8ec9-48f2-ad93-d68d17261cf4 · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5d76cf4-66d7-428d-b646-4fc822d144a0 · inbound
Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b09bb8b-3d4c-4d57-b297-7c818bb3df68 · inbound
ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 75df2934-0219-4bd2-aefc-b5a00ad87e95 · inbound
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0be25dfd-52ef-48ac-a9fa-a35c7f8ac810 · inbound
Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff6e584a-c301-4845-b460-02cd23924670 · inbound
Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 376d516e-062b-4eae-b132-ebc3db2dad25 · inbound
Defending Against Harmful Supervision Hidden in Benign Samples Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45e0734a-62f9-480b-9689-95425a50aea7 · inbound
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · inbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6271f3ad-1a30-4338-bb64-853b857090af · inbound
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc0aa6b-640a-4bd9-8fbc-554c13c04373 · inbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.