Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:12:00.198263Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2501.02629.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:12:00.198263Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T18:05:35.316487Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T18:05:36.436186Z
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 817d811a-bea9-4fe1-b273-ff1dcdfe163d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9e95b9-c2b2-4e17-8c3c-f595d9c0d4e4 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356aceb4-71db-41d3-b880-1cd1814e0392 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Detecting Language Model Attacks with Perplexity
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15828501-0fed-494c-9fa3-c892699c1fbd · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e16a1ca-33e8-44ff-a6bf-b369f1ca77b5 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef17cfbb-4df1-40e0-8152-4d379e61a309 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Editing Factual Knowledge in Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a3ef82-b907-4768-929f-d1a0b43aae87 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Not All Layers of LLMs Are Necessary During Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c6f668-1b55-442d-9112-f632d59b0cc9 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd03847-c403-4c5c-9031-fa97a68b597e · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Transformer Feed-Forward Layers Are Key-Value Memories
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2263ee-0315-4b3f-bce9-9cbb2f6ab265 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Improving alignment of dialogue agents via targeted human judgements
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22194cb1-e6ea-4170-a41c-3769b162b14e · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense The Unreasonable Ineffectiveness of the Deeper Layers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0553aeac-d3a9-4cf6-9bee-1c1c41c8048d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Query-Based Adversarial Prompt Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd15270f-8537-4644-bf2f-f99df611725c · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd34feb-c71b-4dc9-b5c1-ef275e0aebde · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c66fc9-e32c-46df-b5cc-8929810e5d1d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Mistral 7B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0779ee9c-1b6a-46dc-b905-f7ac00642a34 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Plug-and-Play Adaptation for Continuously-updated QA
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ceef9326-51bc-4f94-bee1-53e504c68cf3 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7cb2709-dc9f-46f6-bee9-565192c4fad0 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca0d8b5-c1e1-4d0b-a2a5-6d13fe1f7823 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense On the Robustness of Editing Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3102223b-5883-4723-ad9e-2fd0e0253ba4 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbb7cd6-774e-4785-814e-42b2d7aca953 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b821143-cf40-4637-91ef-88067a9c947d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30461f3c-299c-4b7b-b36e-83593699ea0f · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1c2d77eb-34e2-4113-8980-e7dd83afa6f6 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b8e90e0-e7b7-4591-b5b1-5fc99e542ced · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Forgetting before Learning: Utilizing Parametric Arithmetic for Knowledge Updating in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e064bc-f242-40fe-8c41-b18b89942d3a · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d6cd602b-3a5b-4c6a-8e75-18a887a53159 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6974c78-c451-4c89-9606-34c198f39c7d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60ed1c7-daaa-4e24-9582-259364f8a016 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6c4054-59cc-4e16-932c-5f4eb776b32d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c69872-25fc-44d8-8909-b931c82a464d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Red Teaming Language Models with Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d1322f-e659-4553-bdd2-4a6727c9e481 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb260a5-b4ad-42fe-acb9-685cd9b14112 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Aligning Large Language Models with Human: A Survey
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d605eed-3151-4592-a9e4-07be3a451d65 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d0d064-c725-415b-9838-b80a30507259 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc528ee-c740-47d2-be32-3db6854f155d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e6614dbe-caeb-41e4-a2db-b7e2b8dc2927 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c396f214-5a25-4f17-8294-9a7f92d9d9d3 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a48cb0e-d279-4f54-b9c8-6f606ad9b48d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Causality Analysis for Evaluating the Security of Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f493bcbd-bf32-4a15-88ab-665e9dff68bc · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429e7db1-9b7e-417f-8894-4e69f54a42f4 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22a627c-6552-4140-b8b4-1d66b26cc76e · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Modifying Memories in Transformer Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd6187b-74d8-4d14-8759-126059ee0d5e · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Is the System Message Really Important to Jailbreaks in Large Language Models?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58340c30-63fc-4b84-85a7-fc8062765266 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense online" 'onlinestring :=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 821e50fa-ce4a-4d15-95d3-6f876f92fde8 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense write newline
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405640b3-808d-4b6a-8593-04648e624cc7 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609cf4a7-5b35-4199-b635-f8175ca136eb · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeb30ab9-e8a8-46d3-ba0a-5864db3ddf45 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4907ea61-c84c-4ea3-8014-89b12d0fb983 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f58f70f-7983-4a1b-a01a-d5f1a40e772f · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87820e7a-a0e3-4bd8-87c6-ad7f1086e605 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbd6292-f52e-4d9e-8ea0-906c150e645a · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4bdcad-85ea-48ec-9658-3a9c69c6e72f · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c5354cd-83ed-47a9-abe2-7d24d4d00b05 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 313787b1-5638-439f-8030-e1bee329838d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dccb1ad-7a00-4758-8163-e94637e9e88d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Certifying LLM Safety against Adversarial Prompting
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769d9811-7730-41c3-8663-5a5c632bc3df · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c0d13d-1cb2-4fee-ab23-32fd97d61e10 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense RAIN: Your Language Models Can Align Themselves without Finetuning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1446adb2-08fd-4e25-8eff-8c5b53259345 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38831abc-7422-472f-b9e5-3ed9e91f43cf · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Continual Learning and Private Unlearning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a972048a-9b7d-4dc7-8159-6995379306e2 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4fed2f62-7737-4914-bba4-146f37c96d43 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5bb156-2976-47d8-87f5-f336cd68bd1d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871466a3-1766-4a9a-a112-e67accee42a5 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b36f8b56-654e-495e-9c91-81465a039516 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b7649f-1476-4a73-8a4c-bf646e6c2ef5 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33fc8aed-559c-4a39-a5e4-964e7b34a414 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa2f11b2-e212-4694-9b34-fb7e08c16973 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98a32bbc-29f3-47b6-b6bb-fe767a83158a · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0383e78-96ff-4219-8b2f-fe014c46f662 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Detoxifying Large Language Models via Knowledge Editing
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83729fab-ed01-4188-8143-48411c326c8d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdc802f1-e24e-41c5-9d3a-d17e9e36c5d7 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbroken: How Does LLM Safety Training Fail?
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ac1aeb8-916b-4f5f-b359-c04df564118e · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3ca14b-fb12-45b5-b492-f91c1014ad53 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292305e5-41b3-4e01-b65b-de621a70054c · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848f4694-f3dd-4a14-89e6-89b4cc3f0fa4 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ce040d5-31ee-4569-87ca-d3b310ac162d · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Unresolved cited work
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f748aa9-3291-4ca5-bbba-c7ebdeaf1169 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Large Language Model Unlearning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2bf9e6-ca15-427c-8341-780a20517969 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54cc01dd-f2fa-4115-a247-eefec3a3e180 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea3b3048-abec-45b6-842e-8574a3f63a75 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a343cc1-97f1-43e0-91af-e32e44dc361a · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Weak-to-Strong Jailbreaking on Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8037181-c287-40cb-9391-0a56fd8b1240 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a760bc7-3b29-4361-b134-ce47eafafdd0 · outbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f617707b-9fa4-4d6c-a559-f94ac80b69f1 · inbound
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.