Pith. sign in

Paper Citation Record · LEDGER

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2310.02949.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02949 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.725601Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.659817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52fff216-cd20-4ab6-bdac-3a7b0be0fd59 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:18:27.647836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:6745e8d2d6c75b6d898573df4c4e8c29b57bdfa6e1b9a86fcac26d4022e6986a

Observation d7a14b8b-4a05-474c-aec7-8cd21659b5b8 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.163266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:e05900cd2597e5067b21c9e750257b5b4ea8985c3f69bdcbfd4591e70635126d

Observation b20987fe-9db7-4022-a645-8d8002f9c10b · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.444563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:7d85383340a74a47ab66add0474e1ae6f4c6221ee1a62b0c7b64935bee9097c1

Observation 1107c248-3eca-47d0-ada5-8b11aaa71381 · inbound

Learning to Ask: When LLM Agents Meet Unclear Instruction cites this paper.

Learning to Ask: When LLM Agents Meet Unclear Instruction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:13:28.074522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-23T21:08:42.276002Z digest=sha256:3274d7934498f8435cd2ac9a29301ebd2bea7c3bf71c46c6ef237c53474de3b0

Observation e7154fa9-a96d-40ee-9d84-d5124b36211b · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.176221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:37113bea1bab45b7e59d7ae3934ff8b6e03ff12f7737fb6e3930ed62820d597a

Observation 97539c8c-b6f4-4227-afd2-b0f9359bf3ae · inbound

Steering Language Model Refusal with Sparse Autoencoders cites this paper.

Steering Language Model Refusal with Sparse Autoencoders Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:45:48.602638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:45:48.602638Z digest=sha256:fcab4e88c324ca73ae126f1b2186d31e5ad6077bf472d9d0bce88bcc5c1ca287

Observation 2bbfcfad-e45c-48b8-8223-07a077794e2d · inbound

RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep? cites this paper.

RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep? Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:18:19.480456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:18:19.480456Z digest=sha256:d11ed70548154b041ca95f63cc3049dccee9607de551756dc91d03ad66d2f669

Observation 26b8fe06-b213-4a31-acf4-60fd49e7fbcf · inbound

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models cites this paper.

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:26:52.629530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:26:52.629530Z digest=sha256:9bf17201cdc4c06f696618b9f6f6c393de9e3758c08c1f002d66e404bf864379

Observation 5c404844-43ab-4d8a-8466-32a1ffca7b90 · inbound

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning cites this paper.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.213034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.213034Z digest=sha256:515f7687f69c8d2d0f85e88cac10e7433775415de40eed3caed22a241897d7a9

Observation 4d06135b-4abe-45b6-9b06-1baf948d5a2f · inbound

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings cites this paper.

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:10.938110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:48:10.938110Z digest=sha256:f9e39359b82c5e65860cdc23e299a3ac00bfca5b8adc9c40f9ddabbf7c72cadb

Observation 303e0b88-975d-4769-a3c8-8d356b9950a2 · inbound

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging cites this paper.

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:18:57.322310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:18:57.322310Z digest=sha256:fd14a40d3e8787c8c7ed5836a9eff481305e0c1bda07a2fd781e6658f061bda0

Observation 35f49797-2e6a-4e80-ad93-790aa4ed03df · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.870948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.870948Z digest=sha256:07b4878e442106f0e2836d92b513a1f758c2c41c58dc45efc2d46481a15276d3

Observation 34195a89-002c-4de1-8b08-2cfc602de11c · inbound

LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language cites this paper.

LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:26:44.965317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:26:44.965317Z digest=sha256:a1b5afd26baa5df68aa8ca09ae362ebd49a53bfe56a28e7188168f24e16a21f5

Observation 81b9a09f-61db-463b-8b6d-d0322382fd37 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.132795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.132795Z digest=sha256:35ea947a6abbeafb28088ec46207e3fd058347de358944fbc0dfc3593ebc2cbb

Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.377658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.377658Z digest=sha256:468e72ab97317cf87bb162549054bc9ae1c55fcbb83f8269f5bf842ceac4bd23

Observation 673dce01-3784-422e-be3e-2367f20e51a7 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 239

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.725601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.725601Z digest=sha256:a1efa9cf4f9b24874af2e1deab18053da847d687751e15d345de25154db980df

Observation 61ff8dc4-fde5-4294-86d1-12f67372e810 · inbound

Latent Adversarial Training Improves the Representation of Refusal cites this paper.

Latent Adversarial Training Improves the Representation of Refusal Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.865408Z digest=sha256:e2df1aad8912ab15a8a2db488d504305b97d4ccd6311060b7ba01fb93187e398

Observation b0439625-cd98-473e-864e-872bba7e7498 · inbound

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation cites this paper.

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:53.136494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:53.136494Z digest=sha256:60fc1551ce5b5e9faddfeac719850550f9a2752c7be44ec6ee4f758c6259d66c

Observation 5e371af3-f895-4cb0-9836-39478243d101 · inbound

Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization cites this paper.

Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.291975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.291975Z digest=sha256:6d8cddc4991858d887eb018eee0557bb790a3767d8603941d8b24eb6fbe80ac2

Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.316238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.316238Z digest=sha256:14f9e2d0a84891d7d594bd73aedebd65c8616038b6bce4785e7454b9f0a6f620

Observation 574b4f9d-c622-4a9a-873a-34bff906c65b · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:41.022429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:41.022429Z digest=sha256:a528c2a0657fe269cfccb4f3d9fe13dc0f8e6842df98c9ee619c75771c133630

Observation 8423a4ca-4720-4e34-9c9a-ca8ba783decd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.826882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.826882Z digest=sha256:14a091b442acce0af327bfc75ca3585ae77d8d83414b7700f7dda5ae73b92283

Observation 301aec01-640b-4ea0-af6f-624097a44f17 · inbound

Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs cites this paper.

Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:30.224427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:09:30.224427Z digest=sha256:293c30993fefd8c2ca8b512463c714de627d763755ce91c4429da18440284a64

Observation 916ad6f6-f012-41d2-8414-a9fbb17f21ee · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.069605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.069605Z digest=sha256:4af4fa37010dde87211dc876a54ad53c78ce2a98b91d146c21aac944978f5c3d

Observation b7964131-1802-4170-9105-c6f4ee58b793 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.032498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.032498Z digest=sha256:4bf9e02cb1bdcdaf4580843b4f75b001383d4c21880905a2ed46becf7e0fac3a

Observation e7c2c932-e692-4384-ad68-e40230d919db · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.960696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.960696Z digest=sha256:40dd33f58e01d6636841d85f1a4d4827e0ee0894642055d00f09f68aa8cf6489

Observation bcfcf42a-f78e-4e05-8033-fde1fe9220d5 · inbound

Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes cites this paper.

Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:59.641049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:59.641049Z digest=sha256:1ea77f99f3256e329ff91646bd48db1f78dd99253a92a33d13a8a7198e5999a5

Observation 6c208afa-7b3a-4326-84bd-828646d8ed1c · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:36.832306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:36.832306Z digest=sha256:22524531e0f993554e5bf80f59541310e7fa8247aedb54f079fbf6a34c6afeee

Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.949136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.949136Z digest=sha256:8aa26ec3d113f54b667e0ea5ecc3436240adacd317b1f6575b55b001cf0cdd5a

Observation d1737b1d-a775-4801-93c3-1501881e726a · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.456783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:db99fd3f1f158c6be8759674bc3cd7ffceee001f8f23d7cc443490e375d3dc3e

Observation 4b329b83-cbf4-4dc4-b7f5-9c8b8ee0aaee · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.882098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.882098Z digest=sha256:8b8208aca8eef6a796426ae0465759c1266b8f551a8b2462910de6b814ed1490

Observation 2f46b1c3-f228-4b2d-b67f-0d8c055eba2c · inbound

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift cites this paper.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.950836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.950836Z digest=sha256:41c07f267cbac712b85470430b0b0024d41f34f8f552e623b5046ca779036582

Observation a9274f2a-95f3-4387-aa55-e4ea777a042f · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.772074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.772074Z digest=sha256:d70666d8b09c719f9ac05790a8f428a2eb6dc78b89c24ff66cee14c50e24c843

Observation a3597325-4084-4808-804f-767418eaaee7 · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.503514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.503514Z digest=sha256:abd803ee0b6d08d7706ba1e54cd7d8910510f4dde9474eac204207fe057325b7

Observation 30a2df69-b7c7-4aa0-bc5f-0db29504e29f · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.200106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:1b7362e7aadcba9ab7f7fb5f9266f1f090bd7ef0d8fc4bcc1174eae51d35d83c

Observation 539c423a-598c-47bd-a544-9f2cdb64d3e4 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.626971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:12eb6e1b7499c6b7738741c42fd94b3318c10416f0e2bfa458993309d50bdf71

Observation 4a20819c-bd6d-4dac-97c7-d940b96d1f11 · inbound

Continual Safety Alignment via Gradient-Based Sample Selection cites this paper.

Continual Safety Alignment via Gradient-Based Sample Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:16:54.723710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:16:53.472918Z digest=sha256:2650bf4ff8d89a939b23e70ea99df8a1907075a5afc44fea111045a9801e63e4

Observation 3982e68c-4587-4aa8-b7ef-d5b29f6f0bb2 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.395425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:09c2e8c1c805320546479b3dcb0104e49d7cc2ac9cdb087505637bfeb577b098

Observation ac357e15-b452-4ffa-9014-d6ea0e5def20 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.500341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:1aca32eca93acd0672251ab1e91ba7c349fdbc1a024b83eb9ea1d99d135ee528

Observation 6c007f23-3612-4364-9134-03e5d47b5c66 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.005373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:8a8fc10aafc6c5adb81848b488695c8a1a8bdb35894e69dd57a4dcae300d527e

Observation 026f3de0-b29c-4319-8c5b-8b46131ae44a · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:49.280147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:7fc74e07064746da6969ebd94ba1343967eb896425c4d60d5404e9d33fd0eca9

Observation 6b05f4bd-9405-42b9-bc68-ec7f5a25c467 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.049638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:5582eb08c0109ad1fca4e7dbf84735d4274edd6f9a4e80cb399298b003754da9

Observation ca6d5dbd-88cc-414b-a563-b525b536752d · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.535416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:8719699eee8642ea810405f30bc9b5fb75c30aa5f16000af28304f1141cad958

Observation 363a98b2-587f-40db-bea0-4c105b6ccaa9 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.065497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:5a5e56e3aa267ab43ec48b76181fcc25cd9716f5d1cc04963b49fe9729924d2a

Observation f2e55edf-6079-4b2a-bbfc-0111bee5acf3 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:08:50.464701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:ac71f48659de72d08732a5d229c4514d1116f59e1d8ca2f463a3d72b6597f13e

Observation 0b84fdc6-85e8-4316-85e1-129d7639b7ab · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.003990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:a9bb8449f1a1032e62d9073396b9fa64e92d86aa74404e32d1536a6d563a3b20

Observation edb54687-7a19-4cea-a398-f22302efe7ad · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T16:04:52.548154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:9fbb377ff2d8f6c003ab48529c02a9212b0ecb1ba0e62a826803dc2ed5db3e4c

Observation 7cacd102-5732-4aed-8bc9-feaef39b7a1c · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.646667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:9c49fc78357a66db2cdebb30b05456559e8bee75cdf08586765767a6af891d4f

Observation 42583a00-abd0-4214-a5eb-7b6dfc4c5c72 · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.660121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:7c43f74e5201752e80a937550c65ab22afb7f9726cf5f6da1831e450987d23e1

Observation 86a45327-cef7-4f8f-bcae-13b1650f33df · inbound

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models cites this paper.

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.777900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T13:48:36.304776Z digest=sha256:d3ebc07c2a370279b1823039b499df4c85146174865b4d0fd42764e25e114f60

Observation d0ffbb66-00a2-4a1e-8735-f2177e4e3f6b · inbound

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization cites this paper.

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:13.494789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T07:41:03.219581Z digest=sha256:6404ace7675b0aaee93a4d4330168e862ff93a1af719bfa99e657de796a1f5ea

Observation bf2da3a5-6d73-4855-ad5b-43ac28e24aad · inbound

CSULoRA: Closest Safe Update Low-Rank Adaptation cites this paper.

CSULoRA: Closest Safe Update Low-Rank Adaptation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.752828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:25:01.052975Z digest=sha256:acc1f48705d738542e2c20bf7460a65f3f05a8c77b080458d41325fda5b3f400

Observation 916ab45f-0e62-42bf-937a-b2dac39d344a · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:46:10.220470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:5f28bccdcc8d16f478af6cb8349a790497aa4217233566f7b641cb07be2c8d41

Observation c1a7d728-736d-4eb8-b6dc-091baa690fe3 · inbound

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models cites this paper.

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:18.023666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:22:01.257829Z digest=sha256:cbe0a65af620561cf871266167b84fe5a85064bea6c37af6569a434947c47fd0

Observation 0d5f89a3-ee3c-4a53-9222-020f6c0cd61a · inbound

Jailbreaking Multimodal Large Language Models using Multi-Clip Video cites this paper.

Jailbreaking Multimodal Large Language Models using Multi-Clip Video Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:36:17.122384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:16:48.957645Z digest=sha256:363132f65a276b92ca90645f578b682b8ec8a84405ec5775fba1a615db88cbe8

Observation f7687dd7-8ec9-48f2-ad93-d68d17261cf4 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.889072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:e0931956ea263780973e2ec1b5178a261451cb3840b4f7f674b5c56e3d64a599

Observation d5d76cf4-66d7-428d-b646-4fc822d144a0 · inbound

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts cites this paper.

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:57:38.392850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:32:18.368158Z digest=sha256:515701d75cb898869fb11c2340012e97e5f11a501f3eaa2922471d485f6b1d54

Observation 6b09bb8b-3d4c-4d57-b297-7c818bb3df68 · inbound

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing cites this paper.

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.364793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T10:02:21.293918Z digest=sha256:59992a9c94ea9669c2924d61a5688457577904ba5affa4b217ac1d76a02ef625

Observation 75df2934-0219-4bd2-aefc-b5a00ad87e95 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.686325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:e612e43f4cb0da3ea86afc393bb5ed6d42c7ed73448200cee35ae5e649d5a583

Observation 0be25dfd-52ef-48ac-a9fa-a35c7f8ac810 · inbound

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection cites this paper.

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:18.893870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T20:40:15.506976Z digest=sha256:1edc7b3cc4a5f55c2c069390287b6ab6f2190d1a1e39d371c42c6b4c5b9d6aac

Observation ff6e584a-c301-4845-b460-02cd23924670 · inbound

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations cites this paper.

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.661267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T10:23:47.981377Z digest=sha256:41ee10c2d5f40851932dcf46dcce6cca63cf3c98945bb95253ff09012bfca8f4

Observation 376d516e-062b-4eae-b132-ebc3db2dad25 · inbound

Defending Against Harmful Supervision Hidden in Benign Samples cites this paper.

Defending Against Harmful Supervision Hidden in Benign Samples Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:24:45.321547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T05:29:04.802064Z digest=sha256:cbeb55a9e35680f2e80cfec158ea3da9f31639401da6b6fa4f0b0be9cbeacdd2

Observation 45e0734a-62f9-480b-9689-95425a50aea7 · inbound

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment cites this paper.

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.623496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T06:20:11.322710Z digest=sha256:a32fc0379c7e4503d3254dbee9ffede83c30002f383e9da36ecb9a4b62c005bd

Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:875d81edd61e1643f3bc07d3fa670c063ffa5435f68a13b0e81d1dbfb914116b

Observation 6271f3ad-1a30-4338-bb64-853b857090af · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.311613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.311613Z digest=sha256:9e8258abd58bffed243f50945e616e2a79d7bf83e6700fb691ec9beb4e1b5bac

Observation 6dc0aa6b-640a-4bd9-8fbc-554c13c04373 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.284016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.284016Z digest=sha256:d899fa17f30881c629a2c82dd79e44de6ff56b9152fced3f34495973e4556168

Observation 933e9d9a-13e5-49af-ab4a-1157bbacc3ec · inbound

Agent Safety Should Be a Runtime Contract cites this paper.

Agent Safety Should Be a Runtime Contract Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:13.984292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:13.984292Z digest=sha256:4253e378c52a1055eecce5f451f531612feeb0e16f64388e0449cfa6bd0f9b26