Pith. sign in

Paper Citation Record · LEDGER

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense

As of 11 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2501.02629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02629 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:12:00.198263Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:05:35.316487Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T18:05:36.436186Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 817d811a-bea9-4fe1-b273-ff1dcdfe163d · outbound

This paper cites GPT-4 Technical Report.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.594541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.594541Z digest=sha256:7cf31b3aa56bcbca974cf4466a70e6d10e8d2094c91f7a6b87610006bf10ec8b

Observation 5a9e95b9-c2b2-4e17-8c3c-f595d9c0d4e4 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.600755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.600755Z digest=sha256:30da1f001f0adb80b9c64f9e5dcd59cb0430e4401fa147e34e3e5d399ba6c952

Observation 356aceb4-71db-41d3-b880-1cd1814e0392 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Detecting Language Model Attacks with Perplexity

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.604783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.604783Z digest=sha256:4bfea81c40fd325ebc4bb331c45238b2641e80fc697fef1dbaf09152e338bda7

Observation 15828501-0fed-494c-9fa3-c892699c1fbd · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.610569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.610569Z digest=sha256:311a41db1919597484e16dc6d4a4a5ed258f48c21be8904c626512f6d09169ed

Observation 5e16a1ca-33e8-44ff-a6bf-b369f1ca77b5 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.634527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.634527Z digest=sha256:d49b9d08b13ddd92635bb8222e514fdf801a82279a6c58ac9f9b3c3721f9f48f

Observation ef17cfbb-4df1-40e0-8152-4d379e61a309 · outbound

This paper cites Editing Factual Knowledge in Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Editing Factual Knowledge in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.639096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.639096Z digest=sha256:c48087f61899f3cc26614c5892002fac8cf5b95590fc894970d277bf15ca4c91

Observation 12a3ef82-b907-4768-929f-d1a0b43aae87 · outbound

This paper cites Not All Layers of LLMs Are Necessary During Inference.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Not All Layers of LLMs Are Necessary During Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.658968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.658968Z digest=sha256:60de8d758b074fe09474d01e63efad2416813a747e84a7315780bc8f19898716

Observation 35c6f668-1b55-442d-9112-f632d59b0cc9 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.667865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.667865Z digest=sha256:4b848e8e3d7941ade8c57d6d2c870ed3b31c5ea880bd827a03e4eee4a65f4ed6

Observation ebd03847-c403-4c5c-9031-fa97a68b597e · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Transformer Feed-Forward Layers Are Key-Value Memories

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.674481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.674481Z digest=sha256:31bb26fbed3de28754b62cc2c23eb1d9a54b84ee6bf8886977f1897666049595

Observation 1d2263ee-0315-4b3f-bce9-9cbb2f6ab265 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Improving alignment of dialogue agents via targeted human judgements

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.681514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.681514Z digest=sha256:0fdd8f6b43ebbff8075de6e251b4f4cb17afccf9ff0d9fdaae9fc676cae6f832

Observation 22194cb1-e6ea-4170-a41c-3769b162b14e · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense The Unreasonable Ineffectiveness of the Deeper Layers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.686596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.686596Z digest=sha256:dc956fff2879901327ef27ea436e115139e6652332a95ffdf089f1d5ffa17206

Observation 0553aeac-d3a9-4cf6-9bee-1c1c41c8048d · outbound

This paper cites Query-Based Adversarial Prompt Generation.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Query-Based Adversarial Prompt Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.692798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.692798Z digest=sha256:14ad3b9dc6d09cb6649d27de9951ff8be1875306e6ed9064998435469092580b

Observation bd15270f-8537-4644-bf2f-f99df611725c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.718374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.718374Z digest=sha256:eb76d37a31553e7b85751fb50af17a0e5c22f9df2fce14c5cae9eb56c65c8c89

Observation 5cd34feb-c71b-4dc9-b5c1-ef275e0aebde · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.730055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.730055Z digest=sha256:6fb39ff1bba34049b0d3b4e07a6c8caaf6c46a321158c458a75aaa95a8c1cfc5

Observation 99c66fc9-e32c-46df-b5cc-8929810e5d1d · outbound

This paper cites Mistral 7B.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.768157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.768157Z digest=sha256:8958d9944583c422000ccb04f8ddeaa3fe28db79fef57ef6d9d318b700aab3ea

Observation 0779ee9c-1b6a-46dc-b905-f7ac00642a34 · outbound

This paper cites Plug-and-Play Adaptation for Continuously-updated QA.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Plug-and-Play Adaptation for Continuously-updated QA

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:12:01.234923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.773806Z digest=sha256:499285ed706ff4e49c936c78110704aa3ab2a3fff249f1ab43141e22731e9c5e

Observation ceef9326-51bc-4f94-bee1-53e504c68cf3 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.790507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.790507Z digest=sha256:402b39300cbeac8a51280c6179ffe04cdb88a52b6b2635cca3e4391a44e2211a

Observation d7cb2709-dc9f-46f6-bee9-565192c4fad0 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.798454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.798454Z digest=sha256:bf22d3661a3c18cf5cb4a5a6582c641f1462c774f19638804c6824766c49bcb9

Observation 0ca0d8b5-c1e1-4d0b-a2a5-6d13fe1f7823 · outbound

This paper cites On the Robustness of Editing Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense On the Robustness of Editing Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:12:01.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.815814Z digest=sha256:7436ad8cd3f6faafc09634c4317270c9595c7f7a11231945f26eb071ef7ce823

Observation 3102223b-5883-4723-ad9e-2fd0e0253ba4 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.820103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.820103Z digest=sha256:55b3177d8c80e42093ab85aecd02784cd8fd3bc0fa8bbe47519d3475beddccca

Observation 6fbb7cd6-774e-4785-814e-42b2d7aca953 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.824423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.824423Z digest=sha256:6a2e9ace348fdd6963e6456bd12b3841e8f3f07ad13e8f7fa0a5f555e7f050bb

Observation 7b821143-cf40-4637-91ef-88067a9c947d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.829094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.829094Z digest=sha256:a2653a7641969830aed19c308e0b93b6e8600e772dfd0a1a6a349a61c9a44031

Observation 30461f3c-299c-4b7b-b36e-83593699ea0f · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:02.003560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.833215Z digest=sha256:decb224903ef9cd118646fabced659747e5f299487a9cce0bf1e55f86c394851

Observation 1c2d77eb-34e2-4113-8980-e7dd83afa6f6 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.976778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.842194Z digest=sha256:4852c208b62234d1e1250854f7aee8d69e669211e767c06d681a207336658d35

Observation 7b8e90e0-e7b7-4591-b5b1-5fc99e542ced · outbound

This paper cites Forgetting before Learning: Utilizing Parametric Arithmetic for Knowledge Updating in Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Forgetting before Learning: Utilizing Parametric Arithmetic for Knowledge Updating in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.846476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.846476Z digest=sha256:32ccda954f3f61a44c8c45ec3eec96232e246dbac0a435d5f01ac159969daeb8

Observation 95e064bc-f242-40fe-8c41-b18b89942d3a · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.956566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.850710Z digest=sha256:bbabd220c3185c775cb70a46e18c096b2ac46a03ee4af7a6f2d49621c69761f0

Observation d6cd602b-3a5b-4c6a-8e75-18a887a53159 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.854924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.854924Z digest=sha256:0b8d697b89c0b9943d5d92308a95c59e160d00727a36ed9cf2a0ce68807cef29

Observation b6974c78-c451-4c89-9606-34c198f39c7d · outbound

This paper cites Future Lens: Anticipating Subsequent Tokens from a Single Hidden State.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Future Lens: Anticipating Subsequent Tokens from a Single Hidden State

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.863949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.863949Z digest=sha256:2d1b96b8b45f378925a71aad20fc285a3f6485a16f77ad2ca51477ba89f89453

Observation f60ed1c7-daaa-4e24-9582-259364f8a016 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.868902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.868902Z digest=sha256:43e7833bafb759c3220cda6742a5daa5be4eb8a106c8d746d386bffa46537941

Observation ca6c4054-59cc-4e16-932c-5f4eb776b32d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.873605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.873605Z digest=sha256:c5a8702b4e69a3ef64ab010ad2576a7fc2e41b923afa82509a06311ceb985027

Observation 04c69872-25fc-44d8-8909-b931c82a464d · outbound

This paper cites Red Teaming Language Models with Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Red Teaming Language Models with Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.877974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.877974Z digest=sha256:e47cfe7382516f3e4f2991156e728bb25effe8edea0a8bdaa246332f14bb8474

Observation a6d1322f-e659-4553-bdd2-4a6727c9e481 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.882281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.882281Z digest=sha256:31561078e3408cedfbada56c29976fec550bcef3e20410d8fe328b98d0eae56c

Observation 8fb260a5-b4ad-42fe-acb9-685cd9b14112 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Aligning Large Language Models with Human: A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.904662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.904662Z digest=sha256:8d073ada6fb0330691a9edeb71876bc48afc02625ba9474d34899d2411447159

Observation 6d605eed-3151-4592-a9e4-07be3a451d65 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.912758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.912758Z digest=sha256:cc079d2d2d9218137ae2ce4af0c0d0c524bb03bbcb40454a56a1a092f3a1ec4b

Observation 12d0d064-c725-415b-9838-b80a30507259 · outbound

This paper cites DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.917002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.917002Z digest=sha256:39b195e75341e649526ea0caeb5c88d19e2b823ea9f64bb03c7b9f545da76e8f

Observation ddc528ee-c740-47d2-be32-3db6854f155d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.884760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:11:59.921213Z digest=sha256:c12c9a45840c972eacf7157ab798faf8545e2c28ec53a04d9e812d17162d5296

Observation e6614dbe-caeb-41e4-a2db-b7e2b8dc2927 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.934274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.934274Z digest=sha256:ad253e879b0f92d2ce76ac92477731cbb993b5c23e2f939659d287fc45b3ba4d

Observation c396f214-5a25-4f17-8294-9a7f92d9d9d3 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.939507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.939507Z digest=sha256:1de74b81c3a9eead1b355d9d3fdf17ab348b9df720a393cb3ad3cb58f0b1795b

Observation 8a48cb0e-d279-4f54-b9c8-6f606ad9b48d · outbound

This paper cites Causality Analysis for Evaluating the Security of Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Causality Analysis for Evaluating the Security of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.943595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.943595Z digest=sha256:d73d4873b0b9e819a7dbc113f44aa7a34da0f7b69c0007becb9508144efac739

Observation f493bcbd-bf32-4a15-88ab-665e9dff68bc · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.948432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.948432Z digest=sha256:d3d75a36a62751020d58b27800b5596fa80fbc13a01785a7dea58f45d5f97dcb

Observation 429e7db1-9b7e-417f-8894-4e69f54a42f4 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.954757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.954757Z digest=sha256:780f95506f30bfeaef4fe66ac7340a35ab382a37c63353005225a3c3dfb38fb1

Observation f22a627c-6552-4140-b8b4-1d66b26cc76e · outbound

This paper cites Modifying Memories in Transformer Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Modifying Memories in Transformer Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.965609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.965609Z digest=sha256:de5cd6b5fc6eb0597c99f9cd94b409ff8aac3a6b06e566d81e3b1d5de1c4739c

Observation ccd6187b-74d8-4d14-8759-126059ee0d5e · outbound

This paper cites Is the System Message Really Important to Jailbreaks in Large Language Models?.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Is the System Message Really Important to Jailbreaks in Large Language Models?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.976519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.976519Z digest=sha256:6a51767aee27fe961dfcaa4f92af7dd65809f1a3dd0d56ac81eba6a658f6b93c

Observation 58340c30-63fc-4b84-85a7-fc8062765266 · outbound

This paper cites online" 'onlinestring :=.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense online" 'onlinestring :=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.981056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.981056Z digest=sha256:2ece198039bc17eb5067587d3ac2c257fcaa585f0145124d8a26cae2f921fdfd

Observation 821e50fa-ce4a-4d15-95d3-6f876f92fde8 · outbound

This paper cites write newline.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.986312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.986312Z digest=sha256:4761d1c44bd86ab25f7ed5b52e6b16ef2b7ca42b3812ce994be923c75e5f575d

Observation 405640b3-808d-4b6a-8593-04648e624cc7 · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.990893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.990893Z digest=sha256:7d5d9d84634dde3cc4917706c75bc888e74d6b9f87934fb12b882da2e6161e8b

Observation 609cf4a7-5b35-4199-b635-f8175ca136eb · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.994933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.994933Z digest=sha256:9c1a977385855703ac4b03c5ccbd7d2bc1a1e96263e565bda4f1e793d34ce5ab

Observation eeb30ab9-e8a8-46d3-ba0a-5864db3ddf45 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.001874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.001874Z digest=sha256:42b422d1692f414d54884dea4324dca943358a3ca8634363e485dac7ce1326c3

Observation 4907ea61-c84c-4ea3-8014-89b12d0fb983 · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.008280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.008280Z digest=sha256:27336000d570c47f6933e1ba6e49ab3522bf6516f466937bcd69b82d61baf82c

Observation 8f58f70f-7983-4a1b-a01a-d5f1a40e772f · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.012228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.012228Z digest=sha256:1018144e7525e35be3009ee05ba89f67d12f17a03341851a8d3fa052c283e0d2

Observation 87820e7a-a0e3-4bd8-87c6-ad7f1086e605 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.018197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.018197Z digest=sha256:8b1d28d8b43be1792926ac40d2cee1bfa8285c25dd67ee443b4d234e78cc74d1

Observation 7cbd6292-f52e-4d9e-8ea0-906c150e645a · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.022595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.022595Z digest=sha256:4ee5b6f8bb5d02bb4949ebb81a289e6e995787c1e226abbcbee8efe6f80d8df8

Observation ea4bdcad-85ea-48ec-9658-3a9c69c6e72f · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.733386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:12:00.029572Z digest=sha256:812db8918192464c4c7a3248889af7488dbb5727219d6204e40fdc4f1f262d8a

Observation 3c5354cd-83ed-47a9-abe2-7d24d4d00b05 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.038446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.038446Z digest=sha256:4a80bc1ab3af99e3d47d04f9382ed85a2873f62a1586cc2bfc46f66bffdb8dc8

Observation 313787b1-5638-439f-8030-e1bee329838d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.043077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.043077Z digest=sha256:673a2cafecf0197605667d070bb8581d1a184192ff156e963447fff091d5261d

Observation 8dccb1ad-7a00-4758-8163-e94637e9e88d · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Certifying LLM Safety against Adversarial Prompting

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.047226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.047226Z digest=sha256:2862f7ef417c80e3b28a6dc18ee73acfebfa827eb62177eecc4b89c636a97f5e

Observation 769d9811-7730-41c3-8663-5a5c632bc3df · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.060734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.060734Z digest=sha256:4799473da700f637a7a4a7f731fda68cc323aaeb9497f9fd26a762b8a5d69769

Observation a8c0d13d-1cb2-4fee-ab23-32fd97d61e10 · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.065764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.065764Z digest=sha256:5e0a8b8817b61ecf9d756292276b731a2fc8ce3feec1ee9f45e98b569e0d579c

Observation 1446adb2-08fd-4e25-8eff-8c5b53259345 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.070815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.070815Z digest=sha256:0d8ade3cc15315c06006406d23c83c8385f75120b32656e2181f272782a6c6c1

Observation 38831abc-7422-472f-b9e5-3ed9e91f43cf · outbound

This paper cites Continual Learning and Private Unlearning.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Continual Learning and Private Unlearning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.074676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.074676Z digest=sha256:97601f3203edaa10190758cbe613af3e46497ca6be8b5996f40cc745f3d56729

Observation a972048a-9b7d-4dc7-8159-6995379306e2 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.682616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:12:00.081177Z digest=sha256:2dca25a8dd38b4f008c8026bd9efdaeedf0d5bf3ceb854aab73e394fa2331c1a

Observation 4fed2f62-7737-4914-bba4-146f37c96d43 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.087620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.087620Z digest=sha256:a1166bd16016a3397f51f83436d7c29077ca135587ffbc4869e3a2c5355cabde

Observation 8e5bb156-2976-47d8-87f5-f336cd68bd1d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.092301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.092301Z digest=sha256:438e687d80f3ca373daff29230a071acafbbc6fdfcd4348d18ff96af7389cdae

Observation 871466a3-1766-4a9a-a112-e67accee42a5 · outbound

This paper cites Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.096522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.096522Z digest=sha256:fe4f57ed6fbeb2df3b25c2d02a508497dc2a97d55b5339fa214e72658a00a260

Observation b36f8b56-654e-495e-9c91-81465a039516 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.100680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.100680Z digest=sha256:9309f2dd7e072af76c1f3db61d557c342eeb8c23ca1507ded7c84079be737fc3

Observation f5b7649f-1476-4a73-8a4c-bf646e6c2ef5 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.638915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:12:00.108615Z digest=sha256:7131122f2ee11d77b2d0c8754d65f7e6d3006ed0153dcd6b681ed8ce98fcd2df

Observation 33fc8aed-559c-4a39-a5e4-964e7b34a414 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.113756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.113756Z digest=sha256:794e5910d522223684ddfcdd1b2ad9fa29ebbe341b10fb931999eb4cdcf47f54

Observation fa2f11b2-e212-4694-9b34-fb7e08c16973 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.625820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:12:00.120086Z digest=sha256:3056ad880ef37d101cc726eb37b704544361ac740cfc8bbcbb8d87227277616c

Observation 98a32bbc-29f3-47b6-b6bb-fe767a83158a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.125110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.125110Z digest=sha256:f63440bc88c0d1b98d953d110383e94d4a4d502d8b088e738c4de031a159882b

Observation c0383e78-96ff-4219-8b2f-fe014c46f662 · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Detoxifying Large Language Models via Knowledge Editing

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.130101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.130101Z digest=sha256:d710016ee3f716f367f2a5d2d8bbee4747e94d9c6b79f378e1ce76d83e30a687

Observation 83729fab-ed01-4188-8143-48411c326c8d · outbound

This paper cites Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.134348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.134348Z digest=sha256:e6f0c55c0de5dec774af0b2fa137b8a9eb364c1ea5e11022e03e3aca919f3982

Observation bdc802f1-e24e-41c5-9d3a-d17e9e36c5d7 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbroken: How Does LLM Safety Training Fail?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.140146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.140146Z digest=sha256:5e64995cdbbc9aef70d10dd868f282f8a2054218ffb6a1374d25722dbad48a62

Observation 4ac1aeb8-916b-4f5f-b359-c04df564118e · outbound

This paper cites Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.144197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.144197Z digest=sha256:8a51f562ffea83cca24ef5d0530e42a52b445618f4d94880058a858e0eb40803

Observation 7a3ca14b-fb12-45b5-b492-f91c1014ad53 · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.149220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.149220Z digest=sha256:2f8de29c241bfab5850a914a715ee8aeb2cec4daf841a32a722e1c49e889e11d

Observation 292305e5-41b3-4e01-b65b-de621a70054c · outbound

This paper cites A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.154510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.154510Z digest=sha256:bdbf54c03f417a45a1a7e4d3542b63268ccf25f7d9ef68c2e583792ebed5e103

Observation 848f4694-f3dd-4a14-89e6-89b4cc3f0fa4 · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:01.599943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:12:00.158627Z digest=sha256:4229516142ed7808b2c81fef62ba4b4cb598e90346f3623e46fc806c7a6edfcc

Observation 2ce040d5-31ee-4569-87ca-d3b310ac162d · outbound

This paper cites an unresolved cited work.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.162774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.162774Z digest=sha256:accc3af114315889e85375f198f4077161083624fd2e51e1f02979a763ac3eec

Observation 4f748aa9-3291-4ca5-bbba-c7ebdeaf1169 · outbound

This paper cites Large Language Model Unlearning.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Large Language Model Unlearning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.167216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.167216Z digest=sha256:b6f161d37df8a5a3877700874930aaf8791bcaff9a6206b1ad82572237137a93

Observation 2b2bf9e6-ca15-427c-8341-780a20517969 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.171521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.171521Z digest=sha256:5a611be996f9658d9a9d1be1530e74503190501558b5f4eeceb422cb3894931a

Observation 54cc01dd-f2fa-4115-a247-eefec3a3e180 · outbound

This paper cites Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.175548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.175548Z digest=sha256:177a7aec2bdf78d07e44ca6eb92f3edf2aecab69bb0b9b529469056155184bc1

Observation ea3b3048-abec-45b6-842e-8574a3f63a75 · outbound

This paper cites Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.181315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.181315Z digest=sha256:2a44fae2c530f35649f6b659ba12a9b9ce9ac28a99c94793fbd819c3942a825e

Observation 9a343cc1-97f1-43e0-91af-e32e44dc361a · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Weak-to-Strong Jailbreaking on Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.187801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.187801Z digest=sha256:edf58eeda674b3491758998981d3871142b2ee4789ea69c415190ab972497e97

Observation a8037181-c287-40cb-9391-0a56fd8b1240 · outbound

This paper cites EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.193445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.193445Z digest=sha256:baa891dea6224ac732313d9f074aab5577d80e1bc9bb2fc9cc864c0fb945fc42

Observation 4a760bc7-3b29-4361-b134-ce47eafafdd0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.198263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.198263Z digest=sha256:324da2aa1842596ac9ec89bc1c8e32d603b0535c24a0a225a411cd9ad76a5b41

Pith citing papers

Observation f617707b-9fa4-4d6c-a559-f94ac80b69f1 · inbound

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks cites this paper.

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-05T18:05:36.550269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-05T18:05:35.316487Z digest=sha256:124c81bd2c85b7ea22b8611ddad294df452d6d950b62d1fcd64bee309348741b