Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:21:58.997491Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2509.06338.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:21:58.997491Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c0c05eef-9022-48b9-bc81-2be2ba0ad8d5 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ef95d3b-167c-42da-b9bd-cc6829dec31a · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0476618d-1137-4ef8-bbec-acaaf3cbbfff · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Refusal in Language Models Is Mediated by a Single Direction
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 007f57b3-905f-4231-93e3-a9829da3619a · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c16cb58-1c18-4637-814f-655aa845de18 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6051bc4-d415-4630-a8a8-6d1264b3241d · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Conference’17, July 2017, Washington, DC, USA Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de473e3f-301b-42e0-a2de-80074899320d · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1fa6dd-89a4-4691-840b-fd0d08cafc5e · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412d3b52-7cfc-40dd-bdde-b38298ed82c0 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Exploiting LLM Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b8fbfe-074b-4a14-bc40-02564e602db2 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c91ebf5-5df8-4145-a9c1-4db7efab60e4 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Sah, and Fathi Am- saad
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 21c3e0e8-c25d-43cd-992c-c96839172633 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4037bb0-48ea-40e9-9565-1e8456f2da7a · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Smoothed Embeddings for Robust Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e254e3-1138-4adc-b695-2056d4002ac1 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Large Language Model Supply Chain: Open Problems From the Security Perspective
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645c7d78-d786-4a06-aedc-f6852660446c · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd0a1b7-d4dd-4971-bf98-d8f32c2881e0 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afef8dbc-a808-4655-b771-164f40d105b6 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Mistral 7B
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5705883a-3555-4e6c-8ef4-f5d491cc574a · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0acb6bc5-0617-43f9-bc2c-0db8dbb3d004 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33adecd4-26ea-48d7-8ba7-c3e5aa51a594 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bcb2cb5-11ab-4f37-bcfe-5a802068d797 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c57b97dc-2c7d-41cc-9fa5-5ff5f8a1e86d · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift The Llama 3 Herd of Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1badeca7-5c3a-4086-a07d-8f972f185b85 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0fa113-6fd8-4fec-adaf-dbaa844fbbdd · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Training language models to follow instructions with human feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd95d0fa-b4cd-4204-9979-67abe94b847e · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524ab78b-ec1a-42cd-9647-33b2f93548f0 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046dffd4-09e0-4616-a220-e0e89254513e · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00a2928-b60b-4059-8a47-3cf891666862 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb63d71-039d-43af-8ba0-f4c1c92cc05c · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Adversarial Attacks and Defenses in Large Language Models: Old and New Threats
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a90ab51-7588-49ee-bb68-757c5ce8b108 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f5cc48-e92e-41a7-98fd-cd43789d5a55 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Large Language Models Encode Clinical Knowledge
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0af96e4-467b-472e-b8a7-fea8f4cfceae · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Gemma: Open Models Based on Gemini Research and Technology
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636da4fb-a1e8-44fc-be6e-3d4a58a60e30 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4430dee8-9501-403c-9da4-dc6a2721bfdc · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc27c8e0-f3ea-4bb1-b032-a3853e8a88c3 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Efficient Adversarial Training in LLMs with Continuous Attacks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e240e770-5338-499e-848f-3d75150583eb · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9852238c-7794-4406-9ff0-4cb83cbbe2dd · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a67477-6cb5-446d-9ff7-192b91c29983 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Qwen2.5-1M Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f46b1c3-f228-4b2d-b67f-0d8c055eba2c · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbf7cbd-5c61-4316-9acf-37ac60dd11fa · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ad7cd4e-8ace-4db5-9746-3cd4c5e513c3 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2186f8c6-612a-4134-8225-1f0e9e65fb46 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift When LLMs Meet Cybersecurity: A Systematic Literature Review
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdf2999-273e-4b66-8dbc-9e9a418ae49b · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4a91b5d8-f775-43d0-8c57-2fc15066ee54 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee540167-276e-488f-a6ff-d4a80eda6928 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede4037f-2d00-42b9-9c57-b72ef58b76c2 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0cc4207a-4537-4eec-bd17-af25c1db5170 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a189e524-5468-427c-87ef-59ad64fc4bbd · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66c05a01-f4fc-4b42-8bba-7a527f1e643a · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997b0497-70c9-4b3f-b9f1-d6a0762563d2 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd75bbd-5814-4d9d-bbfe-cba8fde30a34 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Learning and Individual Differences 103 (2023), 102274
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eadeae49-a81d-4dfd-9824-cf9d34b01915 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift In 33rd USENIX Security Symposium (USENIX Security 24)
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5624a3d0-8089-4952-b44e-5374228ece23 · outbound
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.