Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:48.828674Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2505.15753.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:48.828674Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.175611Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T11:37:15.922072Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99b7ab13-ce2b-460a-adf7-242f3bd5d271 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Detecting Language Model Attacks with Perplexity
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb68cd4-3423-4fb7-946d-878e2646fb5f · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64ceb1b-9e7c-49c2-a707-4a12513aec4c · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Qwen technical report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 886a8169-04ba-48e4-b2e7-e55a244fca48 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Constitutional ai: Harmlessness from ai feedback, 2022
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45cab46-1d09-4a98-879c-4da8a4448c56 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safeinfer: Context adaptive decoding time safety alignment for large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98ec7c4c-dc42-4b15-8d62-997bdb9eb3a4 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d460c8a-66b5-425c-ac4e-76febce85e8c · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards the worst-case robustness of large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1712231-c99f-48e3-a1a8-e34fb7c9b811 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Evaluating large language models trained on code, 2021
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3db68c66-159c-4920-bc1a-6c8fe9a743c8 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00503b70-36cd-4723-bc38-02c821a144c7 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safe rlhf: Safe reinforcement learning from human feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 316fc115-9d99-4232-b4e7-4d35eac70880 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Multilingual Jailbreak Challenges in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a248c4bd-a27b-4259-ab97-eee247e8dff8 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1a44d9-31bf-422e-b958-1cae8ab91900 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f58024da-2a64-4c2a-adc5-c54e21b63d16 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Explaining and Harnessing Adversarial Examples
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6df03e-c5fc-4ac7-808a-95433d683301 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5520535c-e1fb-4184-9649-3fd34870b845 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Measuring massive multitask language understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d6f016-d02b-4072-953a-d445c93b68cd · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8206a2c-c27e-4148-b506-cd9e25c144b1 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64e1a99-04dc-4715-9a5d-43ae87f3f4c1 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval AI Alignment: A Comprehensive Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8729fdfa-68d5-4471-96b8-893b157485b5 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1995933e-f583-4470-a0a7-35dbd5bd062b · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31bca02c-52af-4830-80f1-59a011fa3116 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Artprompt: Ascii art-based jailbreak attacks against aligned llms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b50c3f-e2d8-42a3-b948-a0a6529ea65f · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40304df-a0f8-432a-a08a-754a4a82db50 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Dense passage retrieval for open-domain question answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c615d12-22ed-4350-9a2f-985701feaae0 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Buckley, Jason Phang, Samuel R
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5df0e263-d0cd-4aa3-ae20-573e45d7e151 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e0be59-58ec-4d49-8f45-4ca9174e2b49 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a32c2e-917c-475f-8bbb-020ed343ffbf · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c73d0b-3953-49fd-9b60-d83085a05ea3 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards General Text Embeddings with Multi-stage Contrastive Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b586a7-03a1-453f-aca6-8674ddee56fb · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Autodan: Generating stealthy jailbreak prompts on aligned large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417498db-a5c1-4c68-af90-e1fe302d606b · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e2a2a90-9575-48b8-b1c4-1c77ed3fe5db · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a42ad48c-47f4-4623-a68e-88ac9302a7cd · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems , 37:61065–61105, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9dd6051-07a6-4c12-9962-15a75bfd8fe9 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38744573-a46d-4814-aaf5-834a00156e4a · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Position: Adversarial ML for LLMs Is Not Making Any Progress
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e1fa99-f5d6-49d1-923f-23fd1959b7b9 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The probabilistic relevance framework: Bm25 and beyond
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a0ec857-ce8b-41a3-875c-f07ea6f12cb2 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mitigating skeleton key, a new type of generative ai jailbreak technique
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f15bcb96-d1a4-4fa6-8051-2dbd064f51c2 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Fine-tuning mistral 7b large language model for python query response and code generation: A parameter efficient approach
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5bc4c4e-73aa-49d1-bfac-37998a2d099f · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Intriguing properties of neural networks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07420015-32fa-4fbc-a748-10cfb72eab27 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A theoretical understanding of self-correction through in-context alignment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed33eb3c-588d-40f4-b5d4-4f68de34dea2 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Reinforcement Learning for LLM Post-Training: A Survey
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bedce304-5b00-4bfa-a80b-78b18299dbf8 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbroken: How does llm safety training fail? In NeurIPS, 2023
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bdd8198-b371-4ea6-b84c-bdc93100f652 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80923eb6-3302-4e01-aefa-cd38a4aa9859 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Certifiably robust rag against retrieval corruption
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f959c3cd-4489-4b6b-a410-054ac506d400 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Defending chatgpt against jailbreak attack via self-reminders
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2115bb-ba62-431e-b74f-8fc169eb8c44 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7afac52-f6a5-4c8d-ab70-f10d27949b6a · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e0e766-d937-4532-ba55-fbcc41433d8d · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98157b43-ce31-4c47-8cb9-741cf18ad1b1 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The ai alignment problem: why it is hard, and where to start
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4648f707-0efa-4133-9d6f-04b7ed3ff604 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80e84f86-572f-43b4-adf7-e09f88940881 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Boosting jailbreak attack with momentum
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 252c035a-fe45-41e2-80c2-4547273f67ef · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde56ada-4e21-4d97-9e75-fc1a3582ac16 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval On prompt-driven safeguarding for large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dee30cea-29b1-40b1-b489-8b9ad2672dad · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Poisoning Retrieval Corpora by Injecting Adversarial Passages
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4007808-2e3a-4b2b-9c68-e66246ff8219 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Trustworthiness in Retrieval-Augmented Generation Systems: A Survey
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94337f36-ae35-4de6-9ad1-7e7532814868 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d43d4a-9670-4242-8b07-6e026f2d3d14 · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e778b4-da45-4896-8dfd-0f438f6e2eee · outbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e4b5aa-4b40-4323-bfb8-6848d43a9507 · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d7f6248-08d0-45b0-8e87-2bc492581934 · inbound
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6d4da1b-bd9f-48fb-aa91-94f1b4938306 · inbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.