Pith. sign in

Paper Citation Record · LEDGER

Mitigating Jailbreaks with Intent-Aware LLMs

As of 22 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2508.12072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12072 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:10.244362Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T10:00:02.677030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T10:04:51.400037Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c47d5bc-078a-4911-9100-e4c7d9d32834 · outbound

This paper cites Jailbreaking leading safety-aligned LLM s with simple adaptive attacks.

Mitigating Jailbreaks with Intent-Aware LLMs Jailbreaking leading safety-aligned LLM s with simple adaptive attacks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.506187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:09.880216Z digest=sha256:e85ede4f4edf9f1c835b28a83964f7780a2772a6c6474a16e07cc864edd8e292

Observation ff32cf5d-a269-441e-a9f2-ea87259c1f43 · outbound

This paper cites Activating ai safety level 3 protections.

Mitigating Jailbreaks with Intent-Aware LLMs Activating ai safety level 3 protections

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.485732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:09.887473Z digest=sha256:7068ea470fd0ace0e0fe25ca965b80a389ed523e9bd7bf8722c0e75a90fe93c1

Observation 0b033b50-671b-46d4-81b7-46f84c169124 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Mitigating Jailbreaks with Intent-Aware LLMs Refusal in language models is mediated by a single direction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.466366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:09.894162Z digest=sha256:3918ae9f35ec64d766d6cdb533437909b5828344d7774926514b183db812a6ac

Observation ef4bdce6-965c-4458-a52e-d6b2ddea170e · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Mitigating Jailbreaks with Intent-Aware LLMs A General Language Assistant as a Laboratory for Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.902399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.902399Z digest=sha256:fbb03d834f1025593629df795661cc23c4e2216031670cd5f1a97067d4aaf5b5

Observation 3ad2a0fb-09d3-4301-b1c6-f2b2ac4e7ff3 · outbound

This paper cites GASP : Efficient black-box generation of adversarial suffixes for jailbreaking LLM s.

Mitigating Jailbreaks with Intent-Aware LLMs GASP : Efficient black-box generation of adversarial suffixes for jailbreaking LLM s

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.446110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:09.908203Z digest=sha256:ebca5ad8a0524ff3b77a3584018a7d467af0f1e705eecfa6728118424a58977c

Observation 59b184fa-1d4b-4f0d-ada9-f989831c3c4f · outbound

This paper cites Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions.

Mitigating Jailbreaks with Intent-Aware LLMs Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.913375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.913375Z digest=sha256:b406e70e7e95fe8ccbe85778f7c3c2f8c5a532d682ff375386e2984a858c22ab

Observation 7a1a998c-30b8-42e9-b002-bd04ea3bd9c6 · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

Mitigating Jailbreaks with Intent-Aware LLMs Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.919474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.919474Z digest=sha256:4fa688c30f25984610745c4b0a0124cee8fb225e20f1a2698606e617d0f1a1eb

Observation b8bfb122-33d5-4986-b14c-174550b222c9 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Mitigating Jailbreaks with Intent-Aware LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.925144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.925144Z digest=sha256:917e11d1e34dc0b39dfad901ee891305ca3afe45aa403f9014c0a587bbd9254d

Observation 6f0dd45f-00fa-4e13-a78b-cadf03b6f165 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Mitigating Jailbreaks with Intent-Aware LLMs Jailbreaking black box large language models in twenty queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.930800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.930800Z digest=sha256:da47401d545cb549420d62b974ea0877f4e51e506e35878a5a723a7e4f4fad1f

Observation 63de3034-a7a0-40ac-a943-b50d34b6a3f0 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mitigating Jailbreaks with Intent-Aware LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.936340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.936340Z digest=sha256:0eba3d8dd6a8b3de86d84ef2d0b81e455197b0b1bd48939bf847053aec66b0ea

Observation 13d2f2a6-8ea2-4f8a-9396-3cf5042ee6f6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mitigating Jailbreaks with Intent-Aware LLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.943941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.943941Z digest=sha256:651d279850d0ba8c9d02949f324c7b984ef5a3fa1f8f57d76a8207be8b8c7346

Observation d3be2c77-c0f2-4937-92bf-84b0262b70b6 · outbound

This paper cites A mathematical framework for transformer circuits.

Mitigating Jailbreaks with Intent-Aware LLMs A mathematical framework for transformer circuits

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.952336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.952336Z digest=sha256:06ad629e7ee41697a4de058fc36136c39bff16809e369d8d5994e930570ddad7

Observation 3f26b293-9976-48dd-8802-4f36cffd17f4 · outbound

This paper cites The Llama 3 Herd of Models.

Mitigating Jailbreaks with Intent-Aware LLMs The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.958673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.958673Z digest=sha256:6f1a058c2964967f2da0b48aac6a086ca849c55c0debe43939b4f2621c7688cd

Observation 021d5cfe-93c4-4e14-9fa0-793148835295 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Mitigating Jailbreaks with Intent-Aware LLMs Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.965838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.965838Z digest=sha256:aa9bea2cf0e8534531de12bb4b000db1d63400325a37658b8f2564b4156853d2

Observation f0edbdc2-8b21-4269-b1c7-949d6e369a13 · outbound

This paper cites Intention awareness: Improving upon situation awareness in human-centric environments.

Mitigating Jailbreaks with Intent-Aware LLMs Intention awareness: Improving upon situation awareness in human-centric environments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.391025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:09.973500Z digest=sha256:151c2e900dac7e3222fdd44cea1baeee83b62999cf758de9737515cd153f4111

Observation 853b76b3-30c6-4bd7-bd76-5bbaadc4f97c · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Mitigating Jailbreaks with Intent-Aware LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.980047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.980047Z digest=sha256:b66d222ffa73ab2b40c1b2bcbb015f66756783e51a505eac9d1d782b719d4801

Observation 6234d0f2-dcd9-42b5-af94-bb3c41225c08 · outbound

This paper cites Mixtral of Experts.

Mitigating Jailbreaks with Intent-Aware LLMs Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.986599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.986599Z digest=sha256:80696cf699e1965fedaa7019bdb429a0494419dce0c2e90cdd9e19ad82baa101

Observation cc856a12-6e8b-41d8-8f0f-b6485dbe3ec4 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Mitigating Jailbreaks with Intent-Aware LLMs Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.992450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.992450Z digest=sha256:020bb0d03a0af3bf529afb86d9f69f26ea79e2c7fbb1ec535831616eafd92dfc

Observation 44f10ae0-5251-4bde-885a-99f4a207a0b8 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Mitigating Jailbreaks with Intent-Aware LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:09.999175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:09.999175Z digest=sha256:ab17d351087e2cd0e182c83ce430649c13052e62e7a8a5f51e783dfbf8b7b300

Observation 22c555b8-c51d-479d-97ad-bac6f2b5d7b8 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Mitigating Jailbreaks with Intent-Aware LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.005464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.005464Z digest=sha256:e37ff76cbea9b4e57546595c28d3a815a19f75d0e0ad76142480ce06da305fd5

Observation 54677ca1-7d48-4193-af87-dd22cd5eee36 · outbound

This paper cites DeepSeek-V3 Technical Report.

Mitigating Jailbreaks with Intent-Aware LLMs DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.012108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.012108Z digest=sha256:29f083a6f950d19815e6276daa17580e95b57b0f3eaf57a815e9a87d04217a2b

Observation 5338bdcf-806b-413c-a126-cec8fac44fe9 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Mitigating Jailbreaks with Intent-Aware LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.019491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.019491Z digest=sha256:a92a38748578c4b5fbdfb6d06ccd39fa45e5ce6754a9dcf009215b8520eefaf1

Observation 8e860e02-3f08-4eee-8f69-55d3dc824a19 · outbound

This paper cites PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails.

Mitigating Jailbreaks with Intent-Aware LLMs PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.029145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.029145Z digest=sha256:5e7b98b613dd811b82041700ef0bd4a5c6931b682b7bc2401f72f105aa0cbb19

Observation 403c4415-1430-4150-8002-1e9843f09e9d · outbound

This paper cites GPTEval : A survey on assessments of ChatGPT and GPT-4.

Mitigating Jailbreaks with Intent-Aware LLMs GPTEval : A survey on assessments of ChatGPT and GPT-4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.362490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.035174Z digest=sha256:50131c2b48aa5a52ffac60baedde42f6d12238b58be0b8fbbe08abc2c0e7d304

Observation 3ae14a51-dbc4-42b3-90f1-e23a9b1669bf · outbound

This paper cites Interpreting GPT : The logit lens, 2020.

Mitigating Jailbreaks with Intent-Aware LLMs Interpreting GPT : The logit lens, 2020

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.343169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.042023Z digest=sha256:bec6893c84ddf84e64bc180637b7ec558b7ee7f99c26f1f4eeb29568eb3a1558

Observation 03e33d49-f9a8-4592-8c47-fae210e467a2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Mitigating Jailbreaks with Intent-Aware LLMs Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.047010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.047010Z digest=sha256:c58b37636561f918a3aa24f66fcaa53bf06e1c52b4d159b34ba2bde2e86d0b65

Observation fb54500f-058f-4482-bdc6-9d96135326db · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Mitigating Jailbreaks with Intent-Aware LLMs Steering Llama 2 via Contrastive Activation Addition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.052229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.052229Z digest=sha256:dc6f31d51c93a6bfdebd39dba0658f82f2b1c45f1cd651e133be94ef246e416d

Observation 1a225055-0953-4e84-b4e3-ed03311b145e · outbound

This paper cites Rapid Response: Mitigating LLM Jailbreaks with a Few Examples.

Mitigating Jailbreaks with Intent-Aware LLMs Rapid Response: Mitigating LLM Jailbreaks with a Few Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.059063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.059063Z digest=sha256:69861fbd8dc8f47f8a497e4726c23eb16e7068862eeb03cf00beb9a86da9f0e5

Observation 1967755b-1600-452e-a007-171c105495d1 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Mitigating Jailbreaks with Intent-Aware LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.065531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.065531Z digest=sha256:e4c6c07016dcf557461e03979be05ccacfbdc0cc9bdc7114b5f10fa871d13006

Observation 363cc3c7-c03e-409c-b430-ade50ee50753 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Mitigating Jailbreaks with Intent-Aware LLMs Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.071729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.071729Z digest=sha256:6dca10fdca2ff45fe9e0db730e13cf8482bd60b9d5cb87dd2ac4a8ddf0e0e45e

Observation 95a18e53-130b-4ffa-9d6b-9ba0555f5520 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Mitigating Jailbreaks with Intent-Aware LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.078388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.078388Z digest=sha256:bd965f3144fea03d504bc5840b51d469455956f24fd7b818544ecf70d743fbb0

Observation e43708b6-0e92-475d-99dc-534bc8a2f176 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Mitigating Jailbreaks with Intent-Aware LLMs Gpqa: A graduate-level google-proof q&a benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.083742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.083742Z digest=sha256:8ef4fb28bdb5badcb9a68170850dc7ddc7f22d8b4ac75e13aec92277dc527e0d

Observation 4e191eec-8cc5-42bf-866f-9dc69522eafc · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Mitigating Jailbreaks with Intent-Aware LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.089503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.089503Z digest=sha256:4a1d73695a804d76937ace2d3a40f3f02fe3a115ca255e64c1a6154e07a16360

Observation 916aba2d-6834-429c-a140-da9de1f01423 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Mitigating Jailbreaks with Intent-Aware LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.096105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.096105Z digest=sha256:e9cf130957668d0b960ca7b210b7f4a69a494e40bff03f25a5cd61e8d4734123

Observation 75bf6229-c3f8-4972-b8ad-3cf180975230 · outbound

This paper cites A trivial jailbreak against llama 3.

Mitigating Jailbreaks with Intent-Aware LLMs A trivial jailbreak against llama 3

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.284859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.103374Z digest=sha256:4ad27fa67b73c1e6ff9eed3e360e36ca48373d3bb4a2a5fcb5707223d5c873c9

Observation b30ac8b4-1fdd-4b70-ae18-d0e80df57d38 · outbound

This paper cites Hashimoto.

Mitigating Jailbreaks with Intent-Aware LLMs Hashimoto

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.109488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.109488Z digest=sha256:21fc1b2511e189958898b4aeb675f80231afee0b703e30a315cdfdc865868c3b

Observation c8137613-46a6-4763-b7e6-b6edb5f68863 · outbound

This paper cites Steering Language Models With Activation Engineering.

Mitigating Jailbreaks with Intent-Aware LLMs Steering Language Models With Activation Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.115243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.115243Z digest=sha256:3189c39ce02e95766fe5f27fb918add51e8ac51bcd307401f65b138e67eba220

Observation 0f42bf46-0ab5-442e-8e3e-686ad93b263d · outbound

This paper cites Attention is all you need.

Mitigating Jailbreaks with Intent-Aware LLMs Attention is all you need

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.121110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.121110Z digest=sha256:38735fa2c6294ad7a972e8d7f47ee2c9f51bbbd6320c7d6966417b755de8f4ae

Observation e08b5517-3148-4184-8c94-4806f4113af1 · outbound

This paper cites Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment.

Mitigating Jailbreaks with Intent-Aware LLMs Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.238788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.127407Z digest=sha256:e17218bb8cdaa9cf644ddc6d554896031d4b7868ea9cc22d010df391502c8b56

Observation f1670cfe-f217-4d1b-a61c-5d69e4ad09a0 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Mitigating Jailbreaks with Intent-Aware LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.132535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.132535Z digest=sha256:44fd54031733d018829d7dcbb89768a201bfec328b4dbc5288f0ba17ef2b3706

Observation 9443b52e-49d8-4b5e-b54b-912e938cb6bd · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Mitigating Jailbreaks with Intent-Aware LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.137872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.137872Z digest=sha256:86296e61f573f8c386d70fd7e20c4a017aaf3a7570e46b6ec6aee3d8681fddbe

Observation c7396fc9-14c9-4ac8-bf22-f3fa6ae7041f · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Mitigating Jailbreaks with Intent-Aware LLMs Defending chatgpt against jailbreak attack via self-reminders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.143425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.143425Z digest=sha256:5ea5f34af106e2fb522611d55d09710504c86d79d0d1c545377ca6b705e71a03

Observation 008c067d-6be2-406a-86c9-b962e2d64528 · outbound

This paper cites an unresolved cited work.

Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:29:11.197462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.149962Z digest=sha256:417dc90206df5b65760290fc88f1641c43fbeb4d8620227c172d0f75c59d9173

Observation 6bb7a1ac-0db6-4658-973b-29e124550a31 · outbound

This paper cites Qwen3 Technical Report.

Mitigating Jailbreaks with Intent-Aware LLMs Qwen3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.155747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.155747Z digest=sha256:0a8a0ba52b367a50105642d72e93adfa73fb908395e5cc41d6f29998baf8c95f

Observation 21c0f68c-05f5-45de-8f2d-d8090664d7ac · outbound

This paper cites Understanding Refusal in Language Models with Sparse Autoencoders.

Mitigating Jailbreaks with Intent-Aware LLMs Understanding Refusal in Language Models with Sparse Autoencoders

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.162261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.162261Z digest=sha256:c948dcbef19e99560c75a97cb834aacadb167d1a3bbe685dec73e407b06a3c51

Observation a0dae1fc-7da4-42d3-9976-ed7b99e495be · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Mitigating Jailbreaks with Intent-Aware LLMs GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.169290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.169290Z digest=sha256:a793ed40e711b5ddb61b76950132aa8aafc7c0ede7c2ae78414dc87684e4f503

Observation cdaec08b-c9ea-489f-ba3f-ef2362598c0c · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Mitigating Jailbreaks with Intent-Aware LLMs How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.173330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.175799Z digest=sha256:e8c2310ed2194a012befc7bc02d5294b73541e2e19025b10cbfc144cd3ccb901

Observation f8370a41-0248-4bd0-9d69-30bf64fb40aa · outbound

This paper cites Intention Analysis Makes LLMs A Good Jailbreak Defender.

Mitigating Jailbreaks with Intent-Aware LLMs Intention Analysis Makes LLMs A Good Jailbreak Defender

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.184512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.184512Z digest=sha256:c186a911571cd0fdb10cd3ca829c3b7fc7adc87b496fc041dd01d3179e0ed3e9

Observation 91f1dd1e-2f8b-4c24-9cdc-9b38a1ce7133 · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

Mitigating Jailbreaks with Intent-Aware LLMs 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.191260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.191260Z digest=sha256:7b3f2102d30bdb38f474fe6f8af114e5c8da7f27fc31e9643e78a47daa6d6a98

Observation 876c3381-9c27-4c56-a3c4-0a13060e8389 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

Mitigating Jailbreaks with Intent-Aware LLMs Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:29:11.151312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T17:29:10.198323Z digest=sha256:2690d638b8b6051a8cbcc29369387746a808aaca9ee1b535d68275894d8ce417

Observation 5d9f8eab-517e-4fd1-a044-4dd506347ec6 · outbound

This paper cites Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking.

Mitigating Jailbreaks with Intent-Aware LLMs Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.203834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.203834Z digest=sha256:0bfabeca8f4b197020f8107a7a06f631c51c83eb5c589a5fd83b23717d0dd308

Observation ca6244fc-8205-4bb5-9524-e97f91c234ba · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Mitigating Jailbreaks with Intent-Aware LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.214076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.214076Z digest=sha256:1dbcde4d21531f2b64b1d71d8b13c0cc2eaf43a61157af68996f52a04e5f9b56

Observation e39e89d7-4539-43eb-8341-579ba0114310 · outbound

This paper cites write newline.

Mitigating Jailbreaks with Intent-Aware LLMs write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.220344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.220344Z digest=sha256:ff18966f26986295726607ac6988c73065424b2be6afff1b53c05d8e3a674b76

Observation 42bf9968-7716-4cf6-bdbe-b604aa7e73a0 · outbound

This paper cites @esa (Ref.

Mitigating Jailbreaks with Intent-Aware LLMs @esa (Ref

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.228790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.228790Z digest=sha256:b8630ea63b7d1bb7fc9f6f1d46cac4bf5c658eb63edf9eb47381e3f4196735d0

Observation 8055015f-f9de-4743-8421-f05ec6aa6590 · outbound

This paper cites an unresolved cited work.

Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.237302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.237302Z digest=sha256:daf6828e7f0ed5a7b2dc5a958334a6a6cc1e3c2d0b71a9c109187c10c4e5f5c6

Observation ccfee3f8-9b7d-4227-afdb-a4b12c04e79c · outbound

This paper cites an unresolved cited work.

Mitigating Jailbreaks with Intent-Aware LLMs Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.244362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.244362Z digest=sha256:6a4c1ab94806401dacc57e3e500e15ae8fa3895bb3495b28fc628669ef9a95f4

Pith citing papers

Observation 89d30831-1b5d-4210-8782-5bdf095841ab · inbound

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail cites this paper.

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail Mitigating Jailbreaks with Intent-Aware LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T10:04:51.403792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T10:00:02.677030Z digest=sha256:d2a7e8b7df7a676357ec08d597f89e824553e81d122387919b5f3261132ccad0