Pith. sign in

Paper Citation Record · LEDGER

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift

As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2509.06338.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06338 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:21:58.997491Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0c05eef-9022-48b9-bc81-2be2ba0ad8d5 · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:22:00.037107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.723129Z digest=sha256:c0c9a7144cfc036f5891401b74b34ad9b1c9058ee03bfdbffca1d42dc61287ab

Observation 6ef95d3b-167c-42da-b9bd-cc6829dec31a · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.728709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.728709Z digest=sha256:064a7520acfb32d55997fed7c92381bc0174c32d4428d9635b9a7ac3ed1df816

Observation 0476618d-1137-4ef8-bbec-acaaf3cbbfff · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.734558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.734558Z digest=sha256:789c98e0ecf2e150dfd2c5dcc17cd7f65e9096d9311d7062d960d121385dbfb4

Observation 007f57b3-905f-4231-93e3-a9829da3619a · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.739712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.739712Z digest=sha256:0b72fb5350c20d0dc8e249c9891261af19ecbcdde0509763f83d4f909e250714

Observation 4c16cb58-1c18-4637-814f-655aa845de18 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.749481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.749481Z digest=sha256:b3b1c7e861ff74026f0c970f39b42f45aafb9e34c2bb81c825282a44d6d64b59

Observation a6051bc4-d415-4630-a8a8-6d1264b3241d · outbound

This paper cites Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Conference’17, July 2017, Washington, DC, USA Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Conference’17, July 2017, Washington, DC, USA Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:22:00.012051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.754910Z digest=sha256:7745930ec0405766f887c81dd0efced14e888ab80bb8e170f70c90107f8f73b1

Observation de473e3f-301b-42e0-a2de-80074899320d · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.759970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.759970Z digest=sha256:71e30506b38b1c2a10e2aa19fef79b57fe690e504a20ccf7c1881f3934ce38d8

Observation 3a1fa6dd-89a4-4691-840b-fd0d08cafc5e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.765160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.765160Z digest=sha256:6262e020c87a22ab38be07c352371d7c97753bd05f9b2225b8b846b3f8ed42ad

Observation 412d3b52-7cfc-40dd-bdde-b38298ed82c0 · outbound

This paper cites Exploiting LLM Quantization.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Exploiting LLM Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.770502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.770502Z digest=sha256:3812db7b3406c635e4e06cb682b72152bf53503cf423678000a82ffd68bbc99e

Observation f6b8fbfe-074b-4a14-bc40-02564e602db2 · outbound

This paper cites Improved Large Language Model Jailbreak Detection via Pretrained Embeddings.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Improved Large Language Model Jailbreak Detection via Pretrained Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.775054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.775054Z digest=sha256:0f2c365867b24948a94e4241b97890edf1f02ed3f649e7e5e15099dd56ed7154

Observation 4c91ebf5-5df8-4145-a9c1-4db7efab60e4 · outbound

This paper cites Sah, and Fathi Am- saad.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Sah, and Fathi Am- saad

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:21:59.724359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.780267Z digest=sha256:0aaa598c645a9b042a5cc1055049b3a597ed3cb1b9848da23c0e21d74560b68c

Observation 21c3e0e8-c25d-43cd-992c-c96839172633 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.790091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.790091Z digest=sha256:06dc44017840f80c6a24ac6aeb9c4021ec768422f9beaac1af0525751eb652a2

Observation a4037bb0-48ea-40e9-9565-1e8456f2da7a · outbound

This paper cites Smoothed Embeddings for Robust Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Smoothed Embeddings for Robust Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.795021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.795021Z digest=sha256:81eceb1713d1e4fad01f74d4ae5a2c44b43221943e41b6d78611eb172d2b0d3e

Observation 91e254e3-1138-4adc-b695-2056d4002ac1 · outbound

This paper cites Large Language Model Supply Chain: Open Problems From the Security Perspective.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Large Language Model Supply Chain: Open Problems From the Security Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.799851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.799851Z digest=sha256:4d98e60f6b1a53e3338c96fc2f04eb7ba4d14733241b59c35f5926c2b39151fa

Observation 645c7d78-d786-4a06-aedc-f6852660446c · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.804727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.804727Z digest=sha256:490a1dde806c64ab68398c3e899d72f0c5f3ca8798789e229f6edebf9cc34eca

Observation 9bd0a1b7-d4dd-4971-bf98-d8f32c2881e0 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.815294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.815294Z digest=sha256:360d4e3c9c795494adcefe6cca8f8cedd3682da8cca3ae4ebb670525dfd0498b

Observation afef8dbc-a808-4655-b771-164f40d105b6 · outbound

This paper cites Mistral 7B.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.819852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.819852Z digest=sha256:5225a1cbeee0c201db838622e957f86d66cc49635ba401e3ef7e27aa03298bd2

Observation 5705883a-3555-4e6c-8ef4-f5d491cc574a · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.986094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.824746Z digest=sha256:476e256a311cbc3f623397877d4b2cbef3b41c102c6612b95329fcefc97d627b

Observation 0acb6bc5-0617-43f9-bc2c-0db8dbb3d004 · outbound

This paper cites Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.835860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.835860Z digest=sha256:4369ff83a0abe6bc6be81f9bc8b661ef230c802c2b3a8fdf9546dea9b0518580

Observation 33adecd4-26ea-48d7-8ba7-c3e5aa51a594 · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.841537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.841537Z digest=sha256:4607b4b38c65658c5c2e5e06bbeb0b6b969a84b8a7b46f812206a51df6b53868

Observation 7bcb2cb5-11ab-4f37-bcfe-5a802068d797 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.852566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.852566Z digest=sha256:8f28cd60a68513a84d43318b784c6139cfdc51535414fc9d0f809e7642ac0b58

Observation c57b97dc-2c7d-41cc-9fa5-5ff5f8a1e86d · outbound

This paper cites The Llama 3 Herd of Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.857698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.857698Z digest=sha256:dcbfa7f4c31517d5115c0d8fb19602b66d50ef61432b3c6261d74d29f513da1d

Observation 1badeca7-5c3a-4086-a07d-8f972f185b85 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.863345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.863345Z digest=sha256:e8e564a390786a156039a95af0c8feb7c9e0e4d1da8b40fd1c52527235894d19

Observation 8c0fa113-6fd8-4fec-adaf-dbaa844fbbdd · outbound

This paper cites Training language models to follow instructions with human feedback.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Training language models to follow instructions with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.868744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.868744Z digest=sha256:a7ce9f02fbb91238c0987fc02f36ff1807022df37766157a2b5c1d1e0e3019f3

Observation dd95d0fa-b4cd-4204-9979-67abe94b847e · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.873890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.873890Z digest=sha256:44225295a6086081bab9e40e763666edef727a5f1f21b8c34c5f4b6cbc493081

Observation 524ab78b-ec1a-42cd-9647-33b2f93548f0 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.879102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.879102Z digest=sha256:2c28b288884a75d986888ca0034a5cffe1fb17dd23c9f5daf486b27a53321daf

Observation 046dffd4-09e0-4616-a220-e0e89254513e · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.884367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.884367Z digest=sha256:b19f19c0061b6cb30654b920ea002e72605282e43ffff5a2a51548aee263999b

Observation f00a2928-b60b-4059-8a47-3cf891666862 · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.889610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.889610Z digest=sha256:c1add52439dcb25ba65a494bff2d563a320a92822bf0cb10ec6b1e452453bb86

Observation 2fb63d71-039d-43af-8ba0-f4c1c92cc05c · outbound

This paper cites Adversarial Attacks and Defenses in Large Language Models: Old and New Threats.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Adversarial Attacks and Defenses in Large Language Models: Old and New Threats

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.899446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.899446Z digest=sha256:4288961e623339f0d20a6b303affdf12a16bce48392024259393cecb17abbaa1

Observation 4a90ab51-7588-49ee-bb68-757c5ce8b108 · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.905172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.905172Z digest=sha256:38202473efc1eddbdd02b0db97dc3804b17762b145245f595d892087816a1993

Observation 61f5cc48-e92e-41a7-98fd-cd43789d5a55 · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Large Language Models Encode Clinical Knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.910428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.910428Z digest=sha256:d99857cc7d9e348f66d289ff2eef423cbfe3aeed3a460630faa0e9c56a1d11a6

Observation f0af96e4-467b-472e-b8a7-fea8f4cfceae · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Gemma: Open Models Based on Gemini Research and Technology

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.915363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.915363Z digest=sha256:16df93edbf9a919ce27d74300db3f69ea24e44e8a016471acdbd87e0c5e7116b

Observation 636da4fb-a1e8-44fc-be6e-3d4a58a60e30 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.920460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.920460Z digest=sha256:d1ce6020bef33bc4ddb8b4ab9b59dd46efa4748bf3facd9396d804c37dfe9018

Observation 4430dee8-9501-403c-9da4-dc6a2721bfdc · outbound

This paper cites ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.925426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.925426Z digest=sha256:d2cf792217b190045a72ab5b6925e9e455ea5f43271ef952bec50b77c127c2ea

Observation dc27c8e0-f3ea-4bb1-b032-a3853e8a88c3 · outbound

This paper cites Efficient Adversarial Training in LLMs with Continuous Attacks.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Efficient Adversarial Training in LLMs with Continuous Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.931041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.931041Z digest=sha256:7bc01d0eb565f87dad3cc0f30836667b2302a1a11229de757ea7506b4133ea6c

Observation e240e770-5338-499e-848f-3d75150583eb · outbound

This paper cites Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.936322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.936322Z digest=sha256:dd8adc63390de800a8eb846cec9e81ab1dc42937fc5be53f4f0f92e32851da86

Observation 9852238c-7794-4406-9ff0-4cb83cbbe2dd · outbound

This paper cites Qwen3 Technical Report.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.941564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.941564Z digest=sha256:69330d965ca315a3386608526927dcb41fcb886dd9a9bdd73a818ac8827cdc0f

Observation 94a67477-6cb5-446d-9ff7-192b91c29983 · outbound

This paper cites Qwen2.5-1M Technical Report.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Qwen2.5-1M Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.946206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.946206Z digest=sha256:722cd446e448739adea9959e1d6aba0fecb104fed34beeb2ef2c276af11dab0d

Observation 2f46b1c3-f228-4b2d-b67f-0d8c055eba2c · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.950836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.950836Z digest=sha256:3d79421ca47bd792054722fb94ba01a37dafbc36d41d19bfd57fbe152bb5e9e2

Observation bfbf7cbd-5c61-4316-9acf-37ac60dd11fa · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.934929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.955863Z digest=sha256:cffc040906d8547cc2face39aee886370e15531d151de0e760ea7d4496b93e81

Observation 0ad7cd4e-8ace-4db5-9746-3cd4c5e513c3 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.960568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.960568Z digest=sha256:4bfd155def14bcafc91bd6571139dde99ae2c8ea868576196d5cec00949b84c0

Observation 2186f8c6-612a-4134-8225-1f0e9e65fb46 · outbound

This paper cites When LLMs Meet Cybersecurity: A Systematic Literature Review.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift When LLMs Meet Cybersecurity: A Systematic Literature Review

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.966681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.966681Z digest=sha256:23ea9e972683b9d11d1cc3ecfa1e90fe7946acc3c95bef7012f14fe06c1dc1d2

Observation 2bdf2999-273e-4b66-8dbc-9e9a418ae49b · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.919358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.972117Z digest=sha256:e3e7ad55d221fe91154ddc86916c01eac233e8902229d8ccba3b2ed463d7ae67

Observation 4a91b5d8-f775-43d0-8c57-2fc15066ee54 · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.903359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.982439Z digest=sha256:f57ca2a8ebc6dddfb3ca6c677ff706c23ccc8ef0588d38690f54b0d47178ddbe

Observation ee540167-276e-488f-a6ff-d4a80eda6928 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.987023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.987023Z digest=sha256:39c222dcd36254ad77367f878eb06e4e339fcbaa9660e9d56ea81c14281dd595

Observation ede4037f-2d00-42b9-9c57-b72ef58b76c2 · outbound

This paper cites Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T16:21:59.060395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.977344Z digest=sha256:9afc1b8843f23b262372cd3338d1eb2854a7ad73cb188f43880823d82806265c

Observation 0cc4207a-4537-4eec-bd17-af25c1db5170 · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.887498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.992563Z digest=sha256:0f6e42675a2071a4ba6e0cf9e7497af4d54c290a53688fc10d2e4a3e623166f3

Observation a189e524-5468-427c-87ef-59ad64fc4bbd · outbound

This paper cites an unresolved cited work.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:21:59.870957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.997491Z digest=sha256:1adb010c9b883c8dfe358fd97a34a1d571792eb2eaa0e90acd0518c54409f28c

Observation 66c05a01-f4fc-4b42-8bba-7a527f1e643a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.894896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.894896Z digest=sha256:5a40fa29d053f6ed8c39cd3d97db929a8286d4e406ae9fd302ad0a203fce2a5e

Observation 997b0497-70c9-4b3f-b9f1-d6a0762563d2 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.744497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.744497Z digest=sha256:7f0f2dc6d0a00d6223fdecd97d7f69de14f71bca8ec434f33cd17a515636da26

Observation 2cd75bbd-5814-4d9d-bbfe-cba8fde30a34 · outbound

This paper cites Learning and Individual Differences 103 (2023), 102274.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Learning and Individual Differences 103 (2023), 102274

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.831047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.831047Z digest=sha256:6b5e2d9b1a92d281be86c92235eca33d9c292d46d251c1321d83e9e9b1312edf

Observation eadeae49-a81d-4dfd-9824-cf9d34b01915 · outbound

This paper cites In 33rd USENIX Security Symposium (USENIX Security 24).

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift In 33rd USENIX Security Symposium (USENIX Security 24)

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:21:59.960931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T16:21:58.846667Z digest=sha256:2a52ab2bcbc09613c914604d99bc9adb622a0f5fe3199209b607c994053825ad

Observation 5624a3d0-8089-4952-b44e-5374228ece23 · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.809851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.809851Z digest=sha256:45be1cd4dcbf3faf5515480659a51c79b157636961a74ecc70d0d6c18b2f70f4

Pith citing papers

No inbound Pith citation observations are available.