Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:10.244362Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2508.12072.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:10.244362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-08T10:00:02.677030Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T10:04:51.400037Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3c47d5bc-078a-4911-9100-e4c7d9d32834 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Jailbreaking leading safety-aligned LLM s with simple adaptive attacks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff32cf5d-a269-441e-a9f2-ea87259c1f43 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Activating ai safety level 3 protections
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0b033b50-671b-46d4-81b7-46f84c169124 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Refusal in language models is mediated by a single direction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef4bdce6-965c-4458-a52e-d6b2ddea170e · outbound
Mitigating Jailbreaks with Intent-Aware LLMs A General Language Assistant as a Laboratory for Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad2a0fb-09d3-4301-b1c6-f2b2ac4e7ff3 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs GASP : Efficient black-box generation of adversarial suffixes for jailbreaking LLM s
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 59b184fa-1d4b-4f0d-ada9-f989831c3c4f · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1a998c-30b8-42e9-b002-bd04ea3bd9c6 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bfb122-33d5-4986-b14c-174550b222c9 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0dd45f-00fa-4e13-a78b-cadf03b6f165 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Jailbreaking black box large language models in twenty queries
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63de3034-a7a0-40ac-a943-b50d34b6a3f0 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d2f2a6-8ea2-4f8a-9396-3cf5042ee6f6 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3be2c77-c0f2-4937-92bf-84b0262b70b6 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs A mathematical framework for transformer circuits
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f26b293-9976-48dd-8802-4f36cffd17f4 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021d5cfe-93c4-4e14-9fa0-793148835295 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Measuring Massive Multitask Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0edbdc2-8b21-4269-b1c7-949d6e369a13 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Intention awareness: Improving upon situation awareness in human-centric environments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 853b76b3-30c6-4bd7-bd76-5bbaadc4f97c · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6234d0f2-dcd9-42b5-af94-bb3c41225c08 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Mixtral of Experts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc856a12-6e8b-41d8-8f0f-b6485dbe3ec4 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f10ae0-5251-4bde-885a-99f4a207a0b8 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c555b8-c51d-479d-97ad-bac6f2b5d7b8 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54677ca1-7d48-4193-af87-dd22cd5eee36 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs DeepSeek-V3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5338bdcf-806b-413c-a126-cec8fac44fe9 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e860e02-3f08-4eee-8f69-55d3dc824a19 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403c4415-1430-4150-8002-1e9843f09e9d · outbound
Mitigating Jailbreaks with Intent-Aware LLMs GPTEval : A survey on assessments of ChatGPT and GPT-4
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ae14a51-dbc4-42b3-90f1-e23a9b1669bf · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Interpreting GPT : The logit lens, 2020
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 03e33d49-f9a8-4592-8c47-fae210e467a2 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb54500f-058f-4482-bdc6-9d96135326db · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Steering Llama 2 via Contrastive Activation Addition
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a225055-0953-4e84-b4e3-ed03311b145e · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1967755b-1600-452e-a007-171c105495d1 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363cc3c7-c03e-409c-b430-ade50ee50753 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a18e53-130b-4ffa-9d6b-9ba0555f5520 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Direct preference optimization: Your language model is secretly a reward model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43708b6-0e92-475d-99dc-534bc8a2f176 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Gpqa: A graduate-level google-proof q&a benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e191eec-8cc5-42bf-866f-9dc69522eafc · outbound
Mitigating Jailbreaks with Intent-Aware LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916aba2d-6834-429c-a140-da9de1f01423 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75bf6229-c3f8-4972-b8ad-3cf180975230 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs A trivial jailbreak against llama 3
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b30ac8b4-1fdd-4b70-ae18-d0e80df57d38 · outbound
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8137613-46a6-4763-b7e6-b6edb5f68863 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Steering Language Models With Activation Engineering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f42bf46-0ab5-442e-8e3e-686ad93b263d · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Attention is all you need
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08b5517-3148-4184-8c94-4806f4113af1 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f1670cfe-f217-4d1b-a61c-5d69e4ad09a0 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9443b52e-49d8-4b5e-b54b-912e938cb6bd · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7396fc9-14c9-4ac8-bf22-f3fa6ae7041f · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Defending chatgpt against jailbreak attack via self-reminders
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008c067d-6be2-406a-86c9-b962e2d64528 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6bb7a1ac-0db6-4658-973b-29e124550a31 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Qwen3 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c0f68c-05f5-45de-8f2d-d8090664d7ac · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Understanding Refusal in Language Models with Sparse Autoencoders
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0dae1fc-7da4-42d3-9976-ed7b99e495be · outbound
Mitigating Jailbreaks with Intent-Aware LLMs GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdaec08b-c9ea-489f-ba3f-ef2362598c0c · outbound
Mitigating Jailbreaks with Intent-Aware LLMs How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f8370a41-0248-4bd0-9d69-30bf64fb40aa · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Intention Analysis Makes LLMs A Good Jailbreak Defender
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f1dd1e-2f8b-4c24-9cdc-9b38a1ce7133 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876c3381-9c27-4c56-a3c4-0a13060e8389 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Improved few-shot jailbreaking can circumvent aligned language models and their defenses
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d9f8eab-517e-4fd1-a044-4dd506347ec6 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6244fc-8205-4bb5-9524-e97f91c234ba · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39e89d7-4539-43eb-8341-579ba0114310 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs write newline
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42bf9968-7716-4cf6-bdbe-b604aa7e73a0 · outbound
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8055015f-f9de-4743-8421-f05ec6aa6590 · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfee3f8-9b7d-4227-afdb-a4b12c04e79c · outbound
Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d30831-1b5d-4210-8782-5bdf095841ab · inbound
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail Mitigating Jailbreaks with Intent-Aware LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.