Pith. sign in

Paper Citation Record · LEDGER

Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2402.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09063 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:29:16.418332Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8e871a4-f35d-4fcd-8481-0b67b5c75348 · inbound

Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting cites this paper.

Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:18:55.053272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:18:55.053272Z digest=sha256:ea603556bcd442bb58942f954c07bf68b0644003befa878f63ff3caf227503c3

Observation b4c474e4-8e0d-447d-85eb-a6dc4df3e7a4 · inbound

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs cites this paper.

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:19.036502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:19.036502Z digest=sha256:0565b2ccf471c7af665cf712e56b223eca3fc58d3d34ea762e50d349e94b282e

Observation ad40a32b-3f19-4bfc-b54c-e45792a8415a · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.304017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.304017Z digest=sha256:eba4949833123afc4bd8354bced858eb3bace7b38a0a0afab0bee706a3015489

Observation fc9c4729-bf71-48df-b815-87370bc68e8b · inbound

Fast Proxies for LLM Robustness Evaluation cites this paper.

Fast Proxies for LLM Robustness Evaluation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:32:24.110160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:32:24.110160Z digest=sha256:938b85b03efb8030a3b81b07a18dcb66ca3323fe08537943dffe2b9a0d87a6d1

Observation cc335f10-b94a-4214-9dc5-b34398eefa78 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.349571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:3013ff8fd17cf04ef1895d74d5c2ad07de4c232f0b92656aacfd0978ebc24c3b

Observation 28be6349-f425-497c-ad24-49d4f9380f25 · inbound

Adversarial Attack on Large Language Models using Exponentiated Gradient Descent cites this paper.

Adversarial Attack on Large Language Models using Exponentiated Gradient Descent Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:29:16.418332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:29:16.418332Z digest=sha256:9fdadcc0880436eea934f37c1b7bdfa0ec7099ec9bd13632e772ae993f156748

Observation e8e073b0-f459-4c27-a790-27bc91acf10b · inbound

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search cites this paper.

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.716762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.716762Z digest=sha256:a17e1847cfa47693050a5dc7305f49646037f213b233bc11a191a40234235794

Observation 762746ef-a22b-4ab4-90c5-efad89e3f8b1 · inbound

NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation cites this paper.

NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:56:13.194987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:56:13.194987Z digest=sha256:bedc766842f2e3a55be313df69c63fd2d878e4df7d229d0ed16032d63eb2be0a

Observation d96455b2-84ed-4436-8021-783c7711c8d1 · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:25.052460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:25.052460Z digest=sha256:d9eccb4806d5e15521f892f3d8677a536ee8f8d81c5b5c2ba2f48323985375a6

Observation ba850a58-8388-4051-a10f-7b95ce1f8678 · inbound

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs cites this paper.

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.129178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.129178Z digest=sha256:845760c73083623714f69221ea028dc19602575beaa1dfcaa7232b8b9267042c

Observation c0fae038-732b-47ca-9473-9237fc0effa6 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.718850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.718850Z digest=sha256:5a28c1d2358139c6904896deaf2ef8e813dea42b40b539fa3e7a2470bc7352c2

Observation 25011324-4e75-4273-8cdd-0c3901a57082 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.155433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.155433Z digest=sha256:b30ef7aa5f5dc6caadc8445eccf0a4e61fad099eedf45361e06e7ce74baa5e37

Observation 4a90ab51-7588-49ee-bb68-757c5ce8b108 · inbound

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift cites this paper.

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:21:58.905172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:21:58.905172Z digest=sha256:38202473efc1eddbdd02b0db97dc3804b17762b145245f595d892087816a1993

Observation f863855c-2927-461e-94ff-d7ea1978ab0e · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.105707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:6761e32859c1702248d00500d349ca4f03527aee8ed6e133ebceb1293d9a669c

Observation aeb00d55-7c17-4338-9bac-3b1bd8a0f3d2 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:21.589235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:df2da50fc52d51b7be36cd9a4886ace20136fcd3ea238d3eb3f00bcf30945b02

Observation d3268f6a-4563-4f02-bc30-06e7152dc068 · inbound

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes cites this paper.

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:40:36.714300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:40:36.714300Z digest=sha256:704d2412bfaf2373be42e409a19d3e0a463475551a612e6156d70b85133a2e28