Pith. sign in

Paper Citation Record · LEDGER

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08020 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.782513Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved62
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e65c18f2-faa2-4d3d-b5b0-2b7092dfa637 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.240046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.240046Z digest=sha256:2b5d3113547a9698b79ce018f86f744dfed42c9c9ede649342a3060852f41a24

Observation 7a202d84-0c97-4bb1-b685-2d4d6adbcab1 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.325795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.325795Z digest=sha256:d6ae2e066af4398f7315b53b3caf4b353218973f20420d103c88ee6ab16232b6

Observation 444107dd-de74-4159-b95b-a670c2478fb7 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.448232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.448232Z digest=sha256:7694a5980fc1fb66680775da4ca464d39406c9265eb768d9125ff884e9946259

Observation 88d6aace-017c-430c-9f67-184932d69a17 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.767093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.560615Z digest=sha256:1fdf024b3c9821cfb98ddef4a21b66a6346512682e7558521bc195279c109885

Observation 30b9802b-c9f1-4c0e-bfc2-46c455192534 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.617300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.617300Z digest=sha256:f477dfbadbbd73b757c572c1843a7e20c9b2acb54ce968375b5625b9e1d0eecc

Observation 02cb49fa-74e4-4e07-a063-470147ef71b4 · outbound

This paper cites Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.620368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.620368Z digest=sha256:bd9726f19ec63e30bd0a0c9da7a31de24330d623f73abc050761dcc2c20fba32

Observation c8bfc56b-11d7-40ec-ae64-f31441799312 · outbound

This paper cites Pappas, and Eric Wong.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Pappas, and Eric Wong

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.759006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.623324Z digest=sha256:1cd1708a81fd545362793c1c8da418839ddec6aeda29cbfcd7bacf90cd49807d

Observation 6ddbaa36-1a8e-4753-a0cb-f1f1a5097ff2 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.750722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.628809Z digest=sha256:a64db64b55ad338cc64e6449ff90aa65520304c8ce331dafab1e4ddc9b185ed5

Observation 531a0242-9127-4b39-9423-4cd345f196ce · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.631030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.631030Z digest=sha256:90430e079b880227f4da3e831da4d482bf5ca7a55bd30094c646675e805445e2

Observation 476dc5b1-9d65-4071-a5d5-c515f8dd0743 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.633838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.633838Z digest=sha256:bb01441c8b2ee4d26e90c387a6e3e400c44b64c6256cfce0aa82e7d5a6d5a4d2

Observation 4942d2f7-4109-44a6-875d-4337cbeef09b · outbound

This paper cites Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.636474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.636474Z digest=sha256:ca60add5ba6bcbd099e26dc4822d01a61e5931420b44a68619d31382e5da1467

Observation e14f1808-760f-4f9d-9e25-0064345434c5 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.639279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.639279Z digest=sha256:82dc45f258df6cea10724d283b6a2476f8a9311db0174921e632a99f1184ca69

Observation 7cdc94e7-169b-41ab-9e61-63d88ead5892 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.641819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.641819Z digest=sha256:fa34a1b767d706492972527bb3ff8ddd30eba5455f1c2bf8e7454f677ac94d6c

Observation 80b1ca02-e5b2-4eee-9478-02fdb5d7fc13 · outbound

This paper cites Spear Phishing With Large Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Spear Phishing With Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.644775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.644775Z digest=sha256:78ee13b13cbbb025348233c44e7e6f6512a681f0ccd76a3693aa56bbc9e33ba6

Observation 98e29875-b7c1-437f-85a8-100450f892cc · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.647771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.647771Z digest=sha256:d202ad1bfee3ba446fa475fd5cb6d0ef794e8b75fe4ef03425a04c67743d86e2

Observation 9e4f2b11-53a5-4094-8815-5e9e01119140 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.737212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.650243Z digest=sha256:f8e8949dc75dd50b6d8e03fd9295d211d19bf91e17c2d56bbb7cd4aec42c0f76

Observation c6f50e58-8a14-4848-87c0-f017685c0b7d · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.729005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.655580Z digest=sha256:af2b9b9a0822f39c732bff341093f5d5983354108a7917fc4a4d58cb1732fdf3

Observation 87c3da8a-2bd1-4af8-98b4-e89550e2d730 · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.657815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.657815Z digest=sha256:64fd8e123bcb9cf509cc3fc5f71eeaecd4ef6ac2c899965368fb46d3891e8f32

Observation b40b8ec5-7e71-45ec-bedd-8243e5ef1417 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.660236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.660236Z digest=sha256:cb3e71c38008c5508e5d2116918b14e7e79ffe51e107acd2d5bf963314d21b6a

Observation 3b48f360-31bb-43ea-974a-bc0199787763 · outbound

This paper cites Rittichier, and Arjan Durresi.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Rittichier, and Arjan Durresi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.662534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.662534Z digest=sha256:02094e4a10d2604f0cefb339f699f61a707dcd7dba338b98320e6a5f6d09ebf4

Observation e73887ee-08c2-40a3-a43c-aac12429632a · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.665283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.665283Z digest=sha256:3ecd38d7ceca02756c7dd08868d4fc1ea45b45457708c2c69f2649c03d8daa70

Observation 94c0eb51-f047-4d71-8efb-c7589c1f0868 · outbound

This paper cites Model-Editing-Based Jailbreak against Safety-aligned Large Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Model-Editing-Based Jailbreak against Safety-aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.667591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.667591Z digest=sha256:fcee15cf4756cdcdff007844c3ae5d7d8bd1ccd6b29d5cc2acd9dc70900db5c6

Observation 42b7747e-964d-4fe7-b669-fb29a6f538a6 · outbound

This paper cites Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.670577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.670577Z digest=sha256:9d1776644ba6fbbc4daebd8be652568f83bac673635bb074b53f8892076c4ed8

Observation 6db6c64f-ac0e-4da1-992b-425d213778a7 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.673300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.673300Z digest=sha256:84a6d6c2478a072a3056fe600f45dc6b7b3c2c3f7708a6580667e7c71c6cce67

Observation bdf684ec-85ff-47f2-9438-2d466cb5b72c · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:26:14.715425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.676210Z digest=sha256:7ed476241546581735605faf2307f8b92392d89946e926436f30c41dd535b343

Observation 68e302d3-e619-4d0b-84c4-3b73a9b0e679 · outbound

This paper cites The Llama 3 Herd of Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.678386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.678386Z digest=sha256:03f4dc4c976051123658d9d17130224c25fe16139379e471ffb6d6cca2aa1594

Observation 73850077-4b26-41fa-a24e-18ed6395c167 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.707705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.680832Z digest=sha256:9800e64168ef5b8c96d8a6bd4b7f9725710fa2395fa4183e310c96e6fa4ba0a5

Observation 213609da-1b1c-4db5-a740-f181633f182b · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.699821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.683041Z digest=sha256:12eb99d4edc574fd83af26018534fac85ac4212e61caa899de6ba5ac95ef3482

Observation 2c166ecd-6c9c-4812-953d-a714f858e181 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.691654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.685528Z digest=sha256:b825b63103b006158ae198c105009f9f2128094dabaf0ff64b74059c97d5f446

Observation 01691c5f-dae3-40c1-9116-06b610c643ec · outbound

This paper cites Decoding Secret Memorization in Code LLMs Through Token-Level Characterization.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Decoding Secret Memorization in Code LLMs Through Token-Level Characterization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.688117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.688117Z digest=sha256:69751781ef14ef058be3472c538eeac4211873f26b6a576e0dc3a79af63ac813

Observation e3d6dce4-9beb-45f8-9a67-f8cd306c2940 · outbound

This paper cites GPT-4 Technical Report.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.691187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.691187Z digest=sha256:e1f1335422ad77609fa31bdd2eab4b60423db0982a12a16d5786754c5d36eed5

Observation 64cbc338-cbb8-4047-ac27-1d1d234b21d8 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.684118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.693897Z digest=sha256:6c3afedb3cd4b742bb23e72b6e1f92abdf52baf217ed0c59f81d31ef3698a309

Observation 714362e4-ee54-4d44-9a9f-d4e5159de77e · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.676617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.696362Z digest=sha256:6963ca4e3f8e78b7b329a3716489ebf265ddac385dec7481248448909f04d8a5

Observation a3edde7a-2b09-4096-9e4b-c3d793d59ec3 · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.702267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.702267Z digest=sha256:b7433578189bb781133e0f277add8404a1cf3791282f8b10c6d3ff77cb451c8f

Observation 153125d9-9a68-44fb-ae2f-33c5cd59c40e · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.705071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.705071Z digest=sha256:8bf54deea09390c40f14d11152270dff59fceb3b54f48cb040fe94207da6a112

Observation 70402ed4-0f5f-4170-b963-4ad0be384ac7 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.660638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.707736Z digest=sha256:5ebacad03d543ec8be69b12331e9e70ed3dff354e3f5fa58ca6dc799efc3c974

Observation a85a7e06-bc5f-4fb8-90b7-1f3bff35dfed · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.710725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.710725Z digest=sha256:fe23a4be5efbbb06121ce84876b9ac123239a641f010afa5158fda8807875c14

Observation 56e86c98-86db-4e90-9e14-86e48abcbd8e · outbound

This paper cites Adversarial Attacks and Defenses in Large Language Models: Old and New Threats.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Adversarial Attacks and Defenses in Large Language Models: Old and New Threats

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.716377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.716377Z digest=sha256:1ba1ed03714987c9f738957f5ba6125a7c91dfde113ebf013e3e2f061392f25e

Observation c0fae038-732b-47ca-9473-9237fc0effa6 · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.718850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.718850Z digest=sha256:784b7c18eda2616094e400280953edf9dd906b0f9c721701b7cbb0bb0b20e0f6

Observation cfd7c191-22a8-4785-afc3-458679be44a1 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.647814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.721307Z digest=sha256:8e18abc67726f9d480bc558c6ea56cb6e75c3f7a64390dc25611828536881e6e

Observation 36edf4a4-93eb-44fb-bf6a-913a69bd8446 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.723589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.723589Z digest=sha256:799e573b287330107a9d74dc24386634cdcaae20c04cd1a58ae32beff3cddf90

Observation b1c4b421-96e4-4c44-a26b-d25cac4cfb86 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.635425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.725753Z digest=sha256:13f328914ea1ada3e63c0a86c2b6f0371ed8f070c776b6001eaf35b66a65f600

Observation 7c772fc0-2e04-47b0-a9ff-5da75ae40362 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.627959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.728119Z digest=sha256:50b45638eb3c6e3474c28267cc9a91deed5215522720f603779bc88433ae427e

Observation e43774cb-6eaa-4dcc-b7f2-7580eb4a6769 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.730532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.730532Z digest=sha256:cc43b87b737c0763e0aef6b414cce9fecc14637105dcb778c18020bc6b783c62

Observation da060de6-3b05-48f5-8992-3716a41bfc7f · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.733803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.733803Z digest=sha256:4f7649e2ebb7b924f3fc3e0c80f4bfe7627725e8df2edbc63db62cf220604a15

Observation 353eeb47-4bf9-4bcf-b914-8dd221df76fc · outbound

This paper cites Poisoning Language Models During Instruction Tuning.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Language Models During Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.736202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.736202Z digest=sha256:294e400930d076bae59a0ab53ef4be39445c9c3eee274f1568dc989d81f6e106

Observation e13a7e2d-232a-4f9e-9cdb-8fb90bdbf0f4 · outbound

This paper cites Poisoning Attacks against Recommender Systems: A Survey.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Attacks against Recommender Systems: A Survey

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:26:14.151515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.738590Z digest=sha256:0106a342c1c58290a2f2e69e7ba89efe9f99d895135b70980521e8940fe70cb3

Observation 0a8f1628-907c-4c60-a507-6016e1d7eea4 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Finetuned Language Models Are Zero-Shot Learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.741057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.741057Z digest=sha256:0105017b4bd2507324b1d806d759208c3cc3508485f9f5b60627e97cc7e073ad

Observation 78c43e03-bdbb-458f-83aa-5413134e33b9 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.743717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.743717Z digest=sha256:78fc6ddef506266ec3d845cdadc5627cdd663b30410286e3ef7a7ee9f61f8144

Observation f389e0b7-354a-4671-bdfb-2bbe5d8459d3 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.746263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.746263Z digest=sha256:911064556a25d86952dbb8d772ff80c006bae7b46766d3ea8ae2d1d341930343

Observation a4fbffae-b71b-463c-99a2-76741c84795f · outbound

This paper cites Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:26:14.125247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.748695Z digest=sha256:5e108651efc297d44873e6ad35faeff38dd2c2942279530075fa5e5424066a0d

Observation 0e7a2ca2-816b-4295-9c41-9eeadba19a99 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.751199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.751199Z digest=sha256:b741aa7a23fe4918b975e2e5f4686a46ff54d7bf733145e5c1decf81e380a7f1

Observation adf00763-98f3-4885-96fc-cc9efec81ba4 · outbound

This paper cites Qwen2 Technical Report.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.753557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.753557Z digest=sha256:8c7d4903c3a370b68b9f9307687704affd8ee44b2a48555aa37fa565af2fe54b

Observation f7ee0b76-e1ea-426c-b837-ad7777a81516 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Low-Resource Languages Jailbreak GPT-4

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.755908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.755908Z digest=sha256:16f8300f5097e0eefc0d31230fb4c9283c9ce6b3973822e3e2dc6e40b2a652ce

Observation 314cba2b-a856-4c3b-89a3-5eb01feec2ea · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.758637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.758637Z digest=sha256:d056d87c40067f6ab4372568cedb7bdea56fec7e899771b6659f847c891a0233

Observation 041659d5-19be-4c6b-a56d-f0db0a4fc0b7 · outbound

This paper cites When LLMs Meet Cybersecurity: A Systematic Literature Review.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation When LLMs Meet Cybersecurity: A Systematic Literature Review

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.760965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.760965Z digest=sha256:2d0c792d6930bd9f478e6214f9b26a62495f8cced02fa7d99005b2df56793056

Observation adb8111b-9767-47ee-805a-65a144a3ac4c · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.763688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.763688Z digest=sha256:449c4dc33c9c143e31e3c3ae900851e7ae0f324aa27eb2d3dd56c90d75d6611d

Observation 844513cb-37d2-4861-9209-5cd40fa6cf9e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.766148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.766148Z digest=sha256:518903bf9dfbb76db37ac60615cc2db943c67501209d4ddec341206747ef4024

Observation 7dd113a5-6f02-49bb-b354-1774b536c7dc · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.768941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.768941Z digest=sha256:fb52ea13cfebdf560dbd28e354a324d56783852397aaf7137c935068c9cc29de

Observation c7dc684a-6fcc-421b-843b-5c36214d6415 · outbound

This paper cites Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.771379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.771379Z digest=sha256:ae705cfc1e17f4452536d23383f4ec5e2ef7d0c5c9856174f24c6ff228ec4f79

Observation a23e8d93-7f7b-4d19-8c45-61b2b66b355d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.774858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.774858Z digest=sha256:cf84b4fc7e86282e4935a9725f72ac31a3de7a454dd446ba8586b8d64dfff77b

Observation 1681ef50-917c-4518-b035-aac0e19e9642 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.610511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.777653Z digest=sha256:ec1c6c9241b4f43bcdadb2f9f87ee6e89e8543f4a140e26a7d2f8c7d2689cbed

Observation 5bd1987f-b35d-4ec5-bcff-7fa264b0a7be · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.601969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.780234Z digest=sha256:d751b3b0a7fc9aaabc7d5a08e6973f9eab7351cac9517c13b5c68aaf088e4af0

Observation 26d41793-0f18-404f-ae25-3e4c53e3b453 · outbound

This paper cites Hacking is illegal and unethical, and I would never do anything that could put someone’s security at risk.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Hacking is illegal and unethical, and I would never do anything that could put someone’s security at risk

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.593579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.782513Z digest=sha256:390d9558e92db277257ce22530fc9b9a9a3d529559f1aaf69849219424341a65

Observation 8c34e16b-35ee-460b-b07f-cf5d67bd9f90 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.713457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.713457Z digest=sha256:56002c13f13b7d26e37c28ed8ea489008ebefb9eb1b962782d34e3fa174e5725

Observation 61ff9eb1-732c-465e-9779-3fb2647bd335 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.373964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.373964Z digest=sha256:15b3eb7d7e39848b42b79efc02db0cb1150f11e055404c0470ae344b3d880db5

Observation afc78d83-20d1-4167-82b1-cbaad99b9178 · outbound

This paper cites In The Eleventh International Conference on Learning Representations.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation In The Eleventh International Conference on Learning Representations

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.668413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:26:13.699124Z digest=sha256:fbbcbbb0c2286bb089fe978c2699a41a2c8de9e46c07532c34b046f7a2a4c111

Observation f5ba8e02-9b42-4f11-a9f4-f1c4444f7204 · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.652637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.652637Z digest=sha256:72dbc3d813d96b057fb38e078b4f9ea62cd275475f777a5a872e4da7a5be5149

Pith citing papers

No inbound Pith citation observations are available.