Pith. sign in

Paper Citation Record · LEDGER

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2505.15753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15753 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:48.828674Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.175611Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:37:15.922072Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99b7ab13-ce2b-460a-adf7-242f3bd5d271 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.522821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.522821Z digest=sha256:f92b7261f7a7028f796354e6a44bcc73c6a6a177b3d3df1375cb391ca0285a5e

Observation bbb68cd4-3423-4fb7-946d-878e2646fb5f · outbound

This paper cites RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.529201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.529201Z digest=sha256:a0f102edb5287442c6c332c39043b5272f9c991a30006f98413850d57128fe73

Observation c64ceb1b-9e7c-49c2-a707-4a12513aec4c · outbound

This paper cites Qwen technical report.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Qwen technical report

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.609923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.535290Z digest=sha256:67a04233237a446397d5c5c4cdc219d879b2476e8d3838eb3c8e4b2ee069fef0

Observation 886a8169-04ba-48e4-b2e7-e55a244fca48 · outbound

This paper cites Constitutional ai: Harmlessness from ai feedback, 2022.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Constitutional ai: Harmlessness from ai feedback, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.540149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.540149Z digest=sha256:200516fb0bc281ca19d827f7b7d290674e987dde3376e5f59b040ada1acb112e

Observation b45cab46-1d09-4a98-879c-4da8a4448c56 · outbound

This paper cites Safeinfer: Context adaptive decoding time safety alignment for large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safeinfer: Context adaptive decoding time safety alignment for large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.384358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.544775Z digest=sha256:e44b9d5da783d3005b5ede37f6c368f410b4ceb5bc251a5bd5c603653378e675

Observation 98ec7c4c-dc42-4b15-8d62-997bdb9eb3a4 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.550190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.550190Z digest=sha256:5b79e6090c93006fe63866ffd6322e5a24cb888857a69366ede6ba498a4ada16

Observation 5d460c8a-66b5-425c-ac4e-76febce85e8c · outbound

This paper cites Towards the worst-case robustness of large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards the worst-case robustness of large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.555459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.555459Z digest=sha256:8e4608b8a8b1b4d0ae852d7bac387cd274f5fafa7069c74689aa9d33c24bc9e3

Observation d1712231-c99f-48e3-a1a8-e34fb7c9b811 · outbound

This paper cites Evaluating large language models trained on code, 2021.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Evaluating large language models trained on code, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.170794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.559854Z digest=sha256:2b6baa909719a8ef29508bb1c7277b13775315e119019140025b0874b883fc16

Observation 3db68c66-159c-4920-bc1a-6c8fe9a743c8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.565245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.565245Z digest=sha256:e6e49ba0f57cb4709525a2d1d85640193129175a6627b8453cda5755ae2d9dec

Observation 00503b70-36cd-4723-bc38-02c821a144c7 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safe rlhf: Safe reinforcement learning from human feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.667541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.571116Z digest=sha256:ded36dac3987e45bbc4ad533c2aabab3ae4fd77089af99190037d8dd47f595a7

Observation 316fc115-9d99-4232-b4e7-4d35eac70880 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Multilingual Jailbreak Challenges in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.577618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.577618Z digest=sha256:41d81b96cc8616ffee01ca2fb8bc41444b0b6858ead11543047b16f9536bbd75

Observation a248c4bd-a27b-4259-ab97-eee247e8dff8 · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.584004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.584004Z digest=sha256:bcbaeb2ccbd10f1c2f0457123ded2969f1b5f8c4cacd369d0d8a1a73fd2c34b4

Observation 1e1a44d9-31bf-422e-b958-1cae8ab91900 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.589805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.589805Z digest=sha256:0c36d2d77446021f94ffd6f11d3c6ee89af01e9dacee4c3567d077e0b2856a6b

Observation f58024da-2a64-4c2a-adc5-c54e21b63d16 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Explaining and Harnessing Adversarial Examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.594740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.594740Z digest=sha256:20767d9846c5de9c3fcd327e916c83b636d508cc56860e36714416d3845b7fcb

Observation 5c6df03e-c5fc-4ac7-808a-95433d683301 · outbound

This paper cites The Llama 3 Herd of Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.601205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.601205Z digest=sha256:df7396af3a63b843be3610b9b82a38329ceaac86334af47f348043b4b24b92f5

Observation 5520535c-e1fb-4184-9649-3fd34870b845 · outbound

This paper cites Measuring massive multitask language understanding.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.605873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.605873Z digest=sha256:84002f8320f2bc760c06217cb5f47baf342b0162cff1261b438855df03fb8c92

Observation 52d6f016-d02b-4072-953a-d445c93b68cd · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.609882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.609882Z digest=sha256:85f78a11532106c4a3b28619f4d4acb4c8c4d6fa9347499f7f6d1faf589be223

Observation c8206a2c-c27e-4148-b506-cd9e25c144b1 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.615884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.615884Z digest=sha256:97b60a72199192f872603a161d39fc566c6a92d4043bdca85999c14b5a4a2dc1

Observation f64e1a99-04dc-4715-9a5d-43ae87f3f4c1 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval AI Alignment: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.621390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.621390Z digest=sha256:e23a0502707834613d9638e38101462d011bbf0dfb99587b6615841369b514b1

Observation 8729fdfa-68d5-4471-96b8-893b157485b5 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.626254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.626254Z digest=sha256:ff95d8a453f973e1dd3127876763cf818874cc7d1d147ed67bc21ed4a8875a5e

Observation 1995933e-f583-4470-a0a7-35dbd5bd062b · outbound

This paper cites Mistral 7B.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.631589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.631589Z digest=sha256:bf934e0bea82038f32a0d9784a9ae78dbab09b73636fc3ed949186b0f7c3a967

Observation 31bca02c-52af-4830-80f1-59a011fa3116 · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.636714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.636714Z digest=sha256:8c41790888cb55709923389071458791a4d0697c767037af640556815450e29f

Observation 14b50c3f-e2d8-42a3-b948-a0a6529ea65f · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.641406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.641406Z digest=sha256:4adafc944aa932bf5479c1a6566d976736eb351c07ef0eeba78d201f9abace41

Observation b40304df-a0f8-432a-a08a-754a4a82db50 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Dense passage retrieval for open-domain question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.646827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.646827Z digest=sha256:f14f639200613bd99f6c612fcc529e7d4415db7b228c9be6734ef0377c5f3ba5

Observation 1c615d12-22ed-4350-9a2f-985701feaae0 · outbound

This paper cites Buckley, Jason Phang, Samuel R.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Buckley, Jason Phang, Samuel R

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.316907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.651724Z digest=sha256:7375fa9a5ab15e1e7b0c8463a7be73679676c9926192ebc851d7830a32b4f8e6

Observation 5df0e263-d0cd-4aa3-ae20-573e45d7e151 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.657664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.657664Z digest=sha256:75acf65f6953c077dff91d58e83323e5c44ac0d2da9255165569e98fc8e58f49

Observation f0e0be59-58ec-4d49-8f45-4ca9174e2b49 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.663277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.663277Z digest=sha256:73645ad361468c476c37871fbb034d1c1f34736fcd325ada15e95f8ca7f5f2be

Observation 61a32c2e-917c-475f-8bbb-020ed343ffbf · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.669536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.669536Z digest=sha256:5bd365ca718b4e62f74a52c0c255146b6542e2b993a62a47ec1dbcac3841482e

Observation 97c73d0b-3953-49fd-9b60-d83085a05ea3 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.674872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.674872Z digest=sha256:0aaeac44e1ca1529547f65d41a9e59339db258df6c4fceebed949183281a6532

Observation 21b586a7-03a1-453f-aca6-8674ddee56fb · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.679903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.679903Z digest=sha256:f0e8c2d4e419d6f6d34741c5bfac122ea9f6f2fc33f2c108cb480f0ad9b53113

Observation 417498db-a5c1-4c68-af90-e1fe302d606b · outbound

This paper cites Jailbreaking chatgpt via prompt engineering: An empirical study, 2023.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking chatgpt via prompt engineering: An empirical study, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.112343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.685719Z digest=sha256:34b38737b5a83b9443070d52f3012d60c51ab0ae5bff03cf9b190f16759a298a

Observation 3e2a2a90-9575-48b8-b1c4-1c77ed3fe5db · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.690350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.690350Z digest=sha256:ea1c9073e08435b9d7199c6c55c1780c408c4a9347ca29cf6cf3ee237506a268

Observation a42ad48c-47f4-4623-a68e-88ac9302a7cd · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems , 37:61065–61105, 2024.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems , 37:61065–61105, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.015657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.695248Z digest=sha256:d73267ea86a99e86cf36e6ccd3475495859878503ccf3c815df37e7153c6e4e7

Observation f9dd6051-07a6-4c12-9962-15a75bfd8fe9 · outbound

This paper cites Rapid Response: Mitigating LLM Jailbreaks with a Few Examples.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Rapid Response: Mitigating LLM Jailbreaks with a Few Examples

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.700337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.700337Z digest=sha256:8119dea19aee9586922f8848d3adc01ba80f657687713b02a48e399e1b2398fb

Observation 38744573-a46d-4814-aaf5-834a00156e4a · outbound

This paper cites Position: Adversarial ML for LLMs Is Not Making Any Progress.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Position: Adversarial ML for LLMs Is Not Making Any Progress

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.706076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.706076Z digest=sha256:1f1181c75297c87d5048a8f36cc58544020afa23fd0fa232f7f23403e611b83b

Observation 13e1fa99-f5d6-49d1-923f-23fd1959b7b9 · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The probabilistic relevance framework: Bm25 and beyond

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.918748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.711127Z digest=sha256:e7eaf239060dfa2ce0a9e91e9d90e952c264c5182a1a9f1676c91245483c7729

Observation 2a0ec857-ce8b-41a3-875c-f07ea6f12cb2 · outbound

This paper cites Mitigating skeleton key, a new type of generative ai jailbreak technique.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mitigating skeleton key, a new type of generative ai jailbreak technique

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.810370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.716170Z digest=sha256:962c23c5cdba1523623f550613718b00d50f8c4a9f5568d12d55c2fe6bb64baa

Observation f15bcb96-d1a4-4fa6-8051-2dbd064f51c2 · outbound

This paper cites Fine-tuning mistral 7b large language model for python query response and code generation: A parameter efficient approach.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Fine-tuning mistral 7b large language model for python query response and code generation: A parameter efficient approach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.720684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.720684Z digest=sha256:afbaea758543bab2bf82ddab34bd51de7eed65eded526b8a27a118ee37611d06

Observation b5bc4c4e-73aa-49d1-bfac-37998a2d099f · outbound

This paper cites Intriguing properties of neural networks.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Intriguing properties of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.725718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.725718Z digest=sha256:9c96e2f7e6ca3923048e19027bc00d5f96debc0a0d8eecb7d93827c528fed65c

Observation 07420015-32fa-4fbc-a748-10cfb72eab27 · outbound

This paper cites A theoretical understanding of self-correction through in-context alignment.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A theoretical understanding of self-correction through in-context alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.669636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.730652Z digest=sha256:6559ee65a61165c5cf79a1058c00c78b55f64afc185107c2dcd900e490e916a3

Observation ed33eb3c-588d-40f4-b5d4-4f68de34dea2 · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Reinforcement Learning for LLM Post-Training: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.734701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.734701Z digest=sha256:191f570bfc3327c9a6bfaf3b4c77c2266fe0c3ce54c8cf306c454d3aee094d98

Observation bedce304-5b00-4bfa-a80b-78b18299dbf8 · outbound

This paper cites Jailbroken: How does llm safety training fail? In NeurIPS, 2023.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbroken: How does llm safety training fail? In NeurIPS, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.569860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.740043Z digest=sha256:910b50a39ea17cac975ac3289095a70310ceefdebae1fc06bd81b5c271d4e5fa

Observation 9bdd8198-b371-4ea6-b84c-bdc93100f652 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.745729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.745729Z digest=sha256:db8d9f4099e843892e5e818e22b867ffadbc9ed99f7ff006118a866c0ce9b059

Observation 80923eb6-3302-4e01-aefa-cd38a4aa9859 · outbound

This paper cites Certifiably robust rag against retrieval corruption.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Certifiably robust rag against retrieval corruption

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.750950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.750950Z digest=sha256:cde3fdc1bbe44fcb9dad50bed5d40af3c3bf848bf9f053aa90327b7f3dbf104f

Observation f959c3cd-4489-4b6b-a410-054ac506d400 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Defending chatgpt against jailbreak attack via self-reminders

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.756297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.756297Z digest=sha256:bc98ddc32ed2cf97a3ed30536ecdad38cb76a3b7571cb238e2d031be7ca8231b

Observation 3f2115bb-ba62-431e-b74f-8fc169eb8c44 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.435836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.760766Z digest=sha256:40c6f5840331862fa35634f7b224f6778e728c631eb54b19c86e9a1e3473fc48

Observation d7afac52-f6a5-4c8d-ab70-f10d27949b6a · outbound

This paper cites BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.764754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.764754Z digest=sha256:32350501d80849d3b7e09791a9cb36376974f97fc2b61d247a5859219edf14b7

Observation d2e0e766-d937-4532-ba55-fbcc41433d8d · outbound

This paper cites GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.261654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.770152Z digest=sha256:ae9040335d463da2ba4cc6dcfd912d8d954ac194ad404b43066268f179174de2

Observation 98157b43-ce31-4c47-8cb9-741cf18ad1b1 · outbound

This paper cites The ai alignment problem: why it is hard, and where to start.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The ai alignment problem: why it is hard, and where to start

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.774898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.774898Z digest=sha256:028aafe2b638394fe8062b8067075f40009ca279c9cf657c83ed4a1b48e91aaf

Observation 4648f707-0efa-4133-9d6f-04b7ed3ff604 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.030436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.779396Z digest=sha256:c95e5f7642c6c0e110e959a4bb2336e314a8fa4e475969047e0a4736c6339fc1

Observation 80e84f86-572f-43b4-adf7-e09f88940881 · outbound

This paper cites Boosting jailbreak attack with momentum.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Boosting jailbreak attack with momentum

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:49.714847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.785059Z digest=sha256:4a68b64f2a98cca00ed2742138fa0bb0236e805324cb17d36c612d688d8db54c

Observation 252c035a-fe45-41e2-80c2-4547273f67ef · outbound

This paper cites Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.790118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.790118Z digest=sha256:6fd2f872a8de9b91d91376a0b52c23adf4c035551f8a7ed7ba0a07549b513a0f

Observation fde56ada-4e21-4d97-9e75-fc1a3582ac16 · outbound

This paper cites On prompt-driven safeguarding for large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval On prompt-driven safeguarding for large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:49.528671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:48.796220Z digest=sha256:bae34f3adc772e262f1def76f5f2f5cf4c2f6df02bfd3131879eff9385203001

Observation dee30cea-29b1-40b1-b489-8b9ad2672dad · outbound

This paper cites Poisoning Retrieval Corpora by Injecting Adversarial Passages.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Poisoning Retrieval Corpora by Injecting Adversarial Passages

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.802863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.802863Z digest=sha256:c08a55819410c2462edb6792768a69b08414d96e690a83e0e60d963737593a90

Observation e4007808-2e3a-4b2b-9c68-e66246ff8219 · outbound

This paper cites Trustworthiness in Retrieval-Augmented Generation Systems: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.809235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.809235Z digest=sha256:d41d2dc2172652680c531d436d25076e94c083bca3ad1666b3cbefb7a60653e9

Observation 94337f36-ae35-4de6-9ad1-7e7532814868 · outbound

This paper cites ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.816711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.816711Z digest=sha256:ae489786675a5bbfa12a7bee3efc7f29feec2f03af155066cba5a33603a25d76

Observation e7d43d4a-9670-4242-8b07-6e026f2d3d14 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.822543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.822543Z digest=sha256:bbe3903be000055aa7fb0b6eea3b002eba8562b2f5fd66095af01c0ece3ae33c

Observation 26e778b4-da45-4896-8dfd-0f438f6e2eee · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.828674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.828674Z digest=sha256:194061766d98a27cc506f9a35b79e3b795b0292c1bbcefaa4756887b79abb5f1

Pith citing papers

Observation d0e4b5aa-4b40-4323-bfb8-6848d43a9507 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.923828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:26843d12b9ecbdf08498def32bd2da84405cc15408179bb20f5440d48ba1743e

Observation 8d7f6248-08d0-45b0-8e87-2bc492581934 · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:31:19.445142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:6385a714a77969994e630b3875b58de3c738385d64b43b71cf5ca411141902e3

Observation a6d4da1b-bd9f-48fb-aa91-94f1b4938306 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.175611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.175611Z digest=sha256:9b7564192538ed35cb500a8ee05edeba028cafbafafe7368caf86c8abd365b3e