Pith. sign in

Paper Citation Record · LEDGER

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2310.02949.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02949 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:32.316238Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.659817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52fff216-cd20-4ab6-bdac-3a7b0be0fd59 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:18:27.647836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:dbc9ef653a9bdb9dbaf5378786ad7181af4940d327f8e32b3470f2327a79fdf0

Observation d7a14b8b-4a05-474c-aec7-8cd21659b5b8 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.163266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:ca8d4777cde09e435ac942a5b0d1c3d66d12be163c24017f1b470f0cb7ee0a1b

Observation b20987fe-9db7-4022-a645-8d8002f9c10b · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.444563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:565ab3244bc38d1ab30c8ac8ab92a87882b62eafd99a6c979804e7169815553d

Observation 1107c248-3eca-47d0-ada5-8b11aaa71381 · inbound

Learning to Ask: When LLM Agents Meet Unclear Instruction cites this paper.

Learning to Ask: When LLM Agents Meet Unclear Instruction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:13:28.074522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T21:08:42.276002Z digest=sha256:8357a485fb283d1d9f72348397e67f6b5ab383e6eacfda95a4a4313c108c8c61

Observation e7154fa9-a96d-40ee-9d84-d5124b36211b · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.176221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:ffed01f1d4c7f311a01aedbb3f3c9d3c3514f11b9311b2b8b9abdd5bad25ed05

Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.316238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.316238Z digest=sha256:d0371576b71f1d841f2fe704d270bb9b67ebc7dcd7d0792eb5c3d8c715124775

Observation 574b4f9d-c622-4a9a-873a-34bff906c65b · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:41.022429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:41.022429Z digest=sha256:c3a6f9f9bf44373ea97415fda924ed080435ee55dea0c768d58db563589dc6db

Observation 8423a4ca-4720-4e34-9c9a-ca8ba783decd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.826882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.826882Z digest=sha256:06cc873e31c88636ae5f6ab062b5d246f2042d928a18d183cc208629811c89a6

Observation 916ad6f6-f012-41d2-8414-a9fbb17f21ee · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.069605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.069605Z digest=sha256:b97bcaca2949176374e7e98fe07d73306e4a731e1e488d62fac8149fcfb0ccd6

Observation e7c2c932-e692-4384-ad68-e40230d919db · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.960696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.960696Z digest=sha256:3f3ff0618da8d031ddbd58392288d5d6059882f7a14184ab89ae5c9eb8de2cd9

Observation 6c208afa-7b3a-4326-84bd-828646d8ed1c · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:36.832306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:36.832306Z digest=sha256:ab42bec59097c72f68c955196ac5e683a647f0e8f42f42848b42e5c78c746680

Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.949136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.949136Z digest=sha256:18a40fedd92819c576f5efdaf8150fc14493aafebec354301ed1ec3676a1a993

Observation d1737b1d-a775-4801-93c3-1501881e726a · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.456783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:af885f7cce77a1428bf0e62e598bd9b8213d2d014f02e91a571ffe58c48947e7

Observation 4b329b83-cbf4-4dc4-b7f5-9c8b8ee0aaee · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.882098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.882098Z digest=sha256:1613bb23c7f6b91c3b81caee023bf5ee0b2f466286a92d12bb1bb56c66b2e230

Observation a9274f2a-95f3-4387-aa55-e4ea777a042f · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.772074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.772074Z digest=sha256:1d4091b15afc4c33c816b0532281203621f444eee8fbb23624aeb31f0f655922

Observation a3597325-4084-4808-804f-767418eaaee7 · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.503514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.503514Z digest=sha256:736710a0bd7e4d9631b28569962cc77bee03a3737763a3f0bf0f1d3599f832eb

Observation 30a2df69-b7c7-4aa0-bc5f-0db29504e29f · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.200106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:e708134f8b954f27c7464248ba237d905548702c4ee1b30421bf36f55ec046d7

Observation 539c423a-598c-47bd-a544-9f2cdb64d3e4 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.626971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:b2bf7316c137c200990e67042e1e84cb2811060d19e40b0bba6fb35a1949358e

Observation 4a20819c-bd6d-4dac-97c7-d940b96d1f11 · inbound

Continual Safety Alignment via Gradient-Based Sample Selection cites this paper.

Continual Safety Alignment via Gradient-Based Sample Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:16:54.723710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:16:53.472918Z digest=sha256:c551810c4c096cd9f70f22b074e64b29f8ea3cf58359cb22e955b9119febcd05

Observation 3982e68c-4587-4aa8-b7ef-d5b29f6f0bb2 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.395425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:5f5f569901b76f13eacbb4b56623819a4ae8968600ee52eafc3fbe4b8e2db762

Observation ac357e15-b452-4ffa-9014-d6ea0e5def20 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.500341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:e39a1b88258a1a3fe580ffc62cabefdffa7e089b18540f14d0de47f7ac8bb17b

Observation 6c007f23-3612-4364-9134-03e5d47b5c66 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.005373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:4934f2eac88a3f44f9d5c5c0727d03fb504bd404c75cc9093223eb49e51ab44e

Observation 026f3de0-b29c-4319-8c5b-8b46131ae44a · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:49.280147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:33cae8e5cdc0c3afaaa1b803bd005f62587080868334b93e30e82cda44664e5e

Observation 6b05f4bd-9405-42b9-bc68-ec7f5a25c467 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.049638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:c18217c0a22cc79581cdbf43827894428d6a431c0f5dadf2b38446cd21484d90

Observation ca6d5dbd-88cc-414b-a563-b525b536752d · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.535416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:f4f7b3dfc1ac013eac6532a15729b8c46d397043b42543916a9a71e1f366866b

Observation 363a98b2-587f-40db-bea0-4c105b6ccaa9 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.065497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:dd92c9113f00209064f16a64d3933be33983c127ae302ff9106ea89492db71b7

Observation f2e55edf-6079-4b2a-bbfc-0111bee5acf3 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:08:50.464701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:82ff361c820da258a29ba87d01c6b8a28f0c4b665e5023603d4691cb7449198a

Observation 0b84fdc6-85e8-4316-85e1-129d7639b7ab · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.003990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:d24e61728b15968188c6320386f97028b8f7918b4fc8644302ca106749b56cb4

Observation edb54687-7a19-4cea-a398-f22302efe7ad · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T16:04:52.548154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:11584edfed60548822a38d39462fae6905552285bd48499de52c2e8af675f954

Observation 7cacd102-5732-4aed-8bc9-feaef39b7a1c · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.646667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:2264cf9d0bb055afb6a0501aaef649154638bd03127869daf0f51e2279f09e18

Observation 42583a00-abd0-4214-a5eb-7b6dfc4c5c72 · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.660121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:cd8e022d72b4a8ae8c392165f7c86495758677f035f4289784a2a9a76fb56f85

Observation 86a45327-cef7-4f8f-bcae-13b1650f33df · inbound

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models cites this paper.

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.777900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:48:36.304776Z digest=sha256:bed76d09204826b26272e971bc1e1fef0fea26e1297ebcda3536d9079d4d1297

Observation d0ffbb66-00a2-4a1e-8735-f2177e4e3f6b · inbound

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization cites this paper.

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:13.494789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:41:03.219581Z digest=sha256:a1726bc2807a159a14a2fd98fcf0b1df668d9e95fd3793882aabf6a614a00198

Observation bf2da3a5-6d73-4855-ad5b-43ac28e24aad · inbound

CSULoRA: Closest Safe Update Low-Rank Adaptation cites this paper.

CSULoRA: Closest Safe Update Low-Rank Adaptation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.752828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:25:01.052975Z digest=sha256:b001e33ec173055fefc5e63720c695c8cba28414e0f69ea6d3af7471a169589d

Observation 916ab45f-0e62-42bf-937a-b2dac39d344a · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:46:10.220470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:b14084f763dbc9fca43b5411ddf373c1d6f1e6602fbd743161a995e01b6ba62f

Observation c1a7d728-736d-4eb8-b6dc-091baa690fe3 · inbound

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models cites this paper.

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:18.023666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:22:01.257829Z digest=sha256:10c8206c02c334aaa07d9a847ae6ad0591beb14ec853a394a6e5c4a2ea1f9036

Observation 0d5f89a3-ee3c-4a53-9222-020f6c0cd61a · inbound

Jailbreaking Multimodal Large Language Models using Multi-Clip Video cites this paper.

Jailbreaking Multimodal Large Language Models using Multi-Clip Video Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:36:17.122384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:16:48.957645Z digest=sha256:907839687c655f604f77eb29b503db4c84b227abecfe170cd35f2a2644177717

Observation f7687dd7-8ec9-48f2-ad93-d68d17261cf4 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.889072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:4e8e975d6274ee8cc99cba49c2eda8ae12f45cef04f82bbcd6d237b65790d45a

Observation d5d76cf4-66d7-428d-b646-4fc822d144a0 · inbound

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts cites this paper.

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:57:38.392850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:32:18.368158Z digest=sha256:2dd0af7584138faa1e8c49cc9408d44cfd35c48f1bd16f08957e841d5575fc0c

Observation 6b09bb8b-3d4c-4d57-b297-7c818bb3df68 · inbound

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing cites this paper.

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.364793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:02:21.293918Z digest=sha256:16727abf582f1fb8f43b341f536005fc25dc402e9e742cb2f5471e76257c584a

Observation 75df2934-0219-4bd2-aefc-b5a00ad87e95 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.686325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:7faae5c7bbd95674b70ea802bf6a76fa437f38e465a746a9381654b0eb7080b1

Observation 0be25dfd-52ef-48ac-a9fa-a35c7f8ac810 · inbound

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection cites this paper.

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:18.893870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T20:40:15.506976Z digest=sha256:3eb1f9a68d4ac1b7a39931d372ed807fb0ed36450232c9690d25728c663a9da4

Observation ff6e584a-c301-4845-b460-02cd23924670 · inbound

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations cites this paper.

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.661267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T10:23:47.981377Z digest=sha256:bc6786144b99f1c9c174505e1c5e2d12d856e8d41d12a6b181c2323f4f267918

Observation 376d516e-062b-4eae-b132-ebc3db2dad25 · inbound

Defending Against Harmful Supervision Hidden in Benign Samples cites this paper.

Defending Against Harmful Supervision Hidden in Benign Samples Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:24:45.321547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T05:29:04.802064Z digest=sha256:7c14784de1c5968d2a47698e73f8499129f388bdf490ac25707d64c31db66145

Observation 45e0734a-62f9-480b-9689-95425a50aea7 · inbound

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment cites this paper.

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.623496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T06:20:11.322710Z digest=sha256:c0ea77793370e53423dab16983a99a1091fc437d38dc967b62e7de36f09a3ec6

Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:48f59297facfffd09fc2619eeb667ff88cc53fee37356d9b804bf2a1809468dd

Observation 6271f3ad-1a30-4338-bb64-853b857090af · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.311613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.311613Z digest=sha256:4f70d7653999cb0cacf56b0c8c4a8d30b6757476aa499609a8e4ad6b32c26c6c

Observation 6dc0aa6b-640a-4bd9-8fbc-554c13c04373 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.284016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.284016Z digest=sha256:caa7c3dcd1a7248287b7e8e5c9bc0d1d4431dcadb82992aeeab81445ed3c31cf